Deploy failed at Healthcheck with nothing in the logs: long schema setup at boot, and how I moved it to preDeployCommand
adama-diallo-rse
HOBBYOP

23 days ago

Sharing this because it cost me an evening and the symptom is misleading.

Symptom

FastAPI + Postgres app, Dockerfile deploy. A new deployment failed at Network > Healthcheck. No exception, no crash, nothing useful in the deploy logs. Railway kept the previous deployment live, so nothing looked broken, except the new version never went out.

Cause

The app created and updated its schema at startup: ~180 tables through SQLAlchemy create_all, plus a series of idempotent ALTERs. That runs in FastAPI's lifespan, and uvicorn serves nothing until startup is done. So /health wasn't slow, it simply wasn't listening yet. Once startup went past the healthcheck timeout (300s at the time), Railway failed the deploy.

The part that surprised me: I ran uvicorn with --workers 2. Each worker is a full process with its own startup event, so both ran the schema setup at the same time. They don't just double the work, they wait on each other's table locks. That's what pushed startup over the limit.

Fix 1: serialize startup work with a Postgres advisory lock

SELECT pg_try_advisory_lock(:key)

The first process takes the lock and runs the setup. The others poll until it's released, then skip the setup entirely. Startup time no longer depends on the number of workers.

Gotcha #1: derive the key from a stable hash like sha256, not Python's hash(), which is salted per process. Two workers would compute different keys.

Gotcha #2: a session-level advisory lock survives the rollback SQLAlchemy does when returning a connection to the pool. If unlock fails, invalidate() the connection instead of returning it, or the lock lives as long as the process.

Fix 2 (the real one, rolling out now): move it out of startup with preDeployCommand

In railway.json: "deploy": { "preDeployCommand": ["python /app/scripts/migrate_predeploy.py"], "preDeployTimeoutSeconds": 900, "healthcheckPath": "/health" }

The pre-deploy command runs once in the new image, with your service variables and private networking, before the new version starts. If it exits non-zero, the deploy stops and the old version stays live. No container ever serves traffic on a half-applied schema.

In my version, after a successful run the script records RAILWAY_GIT_COMMIT_SHA in a tiny table. On boot, the app looks for its own commit: if the schema was already applied for this code, it skips the setup and only runs a fast drift check (models vs. real schema). If not (local dev, railway up without git), it falls back to the old startup behaviour.

Two things to know about pre-deploy: volumes are not mounted (so don't use it for a SQLite file on a volume), and there's no timeout by default (set preDeployTimeoutSeconds so a stuck migration can't hang the deploy forever).

Takeaways

"Healthcheck failed, nothing in the logs" often means the app wasn't listening yet, not that it crashed.

Every uvicorn/gunicorn worker runs your startup code.

The healthcheck timeout is a ceiling, not a target. If you keep raising it, that work belongs in preDeployCommand.

Private networking isn't available during the build, so migrations belong in pre-deploy or start, never in the build step.

I'll follow up here with real startup times once the pre-deploy version has been running in production for a bit. Happy to answer questions if you're fighting the same thing.

0 Replies

Welcome!

Sign in to your Railway account to join the conversation.

Loading...