5 days ago
Our service backend (project graceful-spontaneity, project ID a3c95a25-766a-4f48-af59-bb47977cc3d9, environment staging, environment ID 1c220600-1df4-433b-8ba2-73909161ed53, service ID 32219559-e1ca-41d8-b0ce-dc92728266b1) has failed to deploy a new commit 7 times in a row over the last hour, always with the same pattern: build succeeds, container starts, Prisma migrations run with no errors ("No pending migrations to apply"), but healthcheck at /api/v1/health/live fails every attempt and the deployment is marked Failed. The currently ACTIVE deployment (older commit) keeps serving traffic and the database normally throughout.
Failed deployment IDs: 6b12d8bb-a3d3-4b5d-8439-d0b63d893d24, 7dae373c-2261-41f5-905c-d031f744768b, cd997223-8748-4137-9968-7e7624856d0d (please also check other attempts around 2026-10-01 10:50-12:00 GMT+3).
What we already checked:
-
Via Railway SSH console, connected to the active container (same image): node dist/main.js starts without errors, and fetch('http://localhost:3000/api/v1/health/live') from inside the container returns status ok immediately.
-
We used the in-dashboard Agent, which diagnosed it as a healthcheck timeout issue and raised healthcheckTimeout from 30s to 60s, then redeployed (cd997223). This did not fix it, still failed, which rules out timeout duration as the cause.
-
Deployment region shown is europe-west4-drams3a, while our configured Scale region is EU West (Amsterdam).
This looks like the same root cause reported by other users here: https://station.railway.com/questions/repeated-healthcheck-failures-on-newly-p-f34d60ff and in a similar thread about "Database proxy unreachable: P1001 errors blocking deploy" - same pattern: new container starts fine, but can't reach the external Postgres database, while the existing active container keeps working against the same database. We're on Neon Postgres, so this doesn't look specific to one DB provider, it looks like a Railway-side networking/egress issue affecting newly provisioned containers reaching external databases.
Could someone from the team look into networking/egress for newly provisioned containers on this service? We're blocked from shipping a bug fix to production.
2 Replies
5 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 5 days ago
2 days ago
New deployment fails
|
vIs the process listening on $PORT?
|+----+----+
| |
No Yes
| |
Fix bind v
(0.0.0.0) Does curl localhost:$PORT/health return 200?
|
+----+----+
| |
No Yes
| |
App issue v
Test with:
Host: healthcheck.railway.app
|
+----+----+
| |
No Yes
| |
Middleware / v
redirect Does the healthcheck request
appear in the app logs?
|
+----+----+
| |
Yes No
| |
Investigate Strong indication
HTTP response of Railway/networking2 days ago
Your migration output contradicts your own diagnosis, and I think that's the most useful thing in the report.
If Prisma ran and reported "No pending migrations to apply" inside the failing container, that container reached Neon successfully. A container that can't egress to an external Postgres fails at migration, not after it. So whatever is happening, it isn't newly provisioned containers being unable to reach the external DB — that's ruled out by your own logs.
The test you ran doesn't test what's failing. You SSH'd into the active container and fetched http://localhost:3000/api/v1/health/live. That confirms the app serves on loopback in a container that's already healthy. Railway's healthcheck doesn't originate inside the container — it arrives over the network with Host: healthcheck.railway.app. An app bound to 127.0.0.1 instead of 0.0.0.0 passes your test and fails the real probe every single time, and nothing you've checked distinguishes the two.
For NestJS specifically: app.listen(port) binds to all interfaces, but app.listen(port, 'localhost') or '127.0.0.1' does not. Check that call, and check that the port comes from process.env.PORT rather than a hardcoded 3000 — Railway assigns PORT and the probe goes to the assigned one.
Two more things that fail exactly this way and that your evidence doesn't exclude:
Host header rejection. If the new commit added a host allowlist, CORS-adjacent middleware, a redirect-to-HTTPS rule, or helmet with strict settings, the probe's Host: healthcheck.railway.app can get a 301/400 instead of 200. Your localhost fetch sends Host: localhost, so it sails through. Worth running this from inside the container to replicate the real request:
curl -sS -o /dev/null -w '%{http_code}\n' \
-H 'Host: healthcheck.railway.app' \
http://127.0.0.1:$PORT/api/v1/health/live
Anything other than 200 is your answer. A 3xx is the giveaway for redirect middleware.
A health endpoint that waits on something slow. If /health/live touches the DB, a cold Neon branch can cold-start for several seconds — but raising the timeout to 60s and still failing argues against this, and points back at bind address or host header.
The one thing in your report nobody has addressed is the region mismatch: deployments landing in europe-west4-drams3a while your configured scale region is EU West (Amsterdam). If your Neon instance has IP allowlisting enabled, a container deployed in a different region egresses from different addresses, and the allowlist blocks it — except migrations succeeded, so that can't be it either. Still worth asking staff why the deploy region differs from the configured one, as a separate question from the healthcheck.
Order I'd go in: check the listen() bind address first, then run the curl above with the real Host header from inside a container. Between them, those cover the large majority of "works on localhost, fails healthcheck" cases, and both are things the Railway team would ask you for anyway before looking at networking.