a month ago
Project: rave (e31f47e4-752f-477e-a3f1-0476a3a37afa), service "web" (e1880986-6983-42e1-a4c4-94a68fafd188), production environment.
Every new deployment of this service has failed healthcheck for the last several hours (5 consecutive attempts today), while the currently active deployment from 2 days ago keeps serving fine. Symptoms:
- Build succeeds, image pushes successfully.
-
- Deploy phase succeeds in ~2-4s.
-
- Container starts: logs "Starting Container", then Next.js prints "Ready in 0ms", then ABSOLUTE SILENCE — zero further log lines — for the entire 5-minute healthcheck window.
-
- Healthcheck path "/" gets "service unavailable" on all 14 attempts, every single time.
-
- Railway's own deployment detail view labels it "Deployment failed during the network process" > "Network > Healthcheck".
-
- Network Logs (HTTP tab) for the failed deployment show zero recorded requests at all — not even failed ones.
What I've already ruled out:
- Not the app code — redeployed the exact same already-built image (same digest) with zero changes: failed identically.
- Not a stale build — triggered a brand-new build from the latest commit on main: failed identically (different image digest, same failure).
- Not memory/CPU — memory usage ~345MB against a 24GB limit, nowhere near the limit.
- Tried the known Railway V2 runtime issue where the app needs to bind to the IPv6 wildcard (::) instead of 0.0.0.0 (found via a few older Central Station threads describing this same symptom) — changed apps/web/Dockerfile's HOSTNAME env from 0.0.0.0 to :: and redeployed: failed identically.
- Compared healthcheckPath/timeout/region/replica config between the currently-active successful deployment and the failing ones via
railway deployment list --json— byte-for-byte identical config. - Built-in "Diagnose" AI feature in the dashboard got stuck on "Fetching deployment info"/"Fetching metrics" both times I tried it and never returned a result.
This really looks like something broken on the platform/infra side specific to spinning up NEW instances of this one service (private networking route not actually being wired up to the healthcheck prober for new containers?), not anything in our code or config. Would appreciate an engineer with backend/infra visibility taking a look at deployment 864041f4-cce3-4fac-b650-3fb644918c59 (or any of today's failed ones) — this is blocking shipping a legally-required Privacy Policy/Terms fix.
5 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
Make sure you have the PORT variable set to the same port your app is listening on, and that the path / returns 200 as the status code.
a month ago
same - can't deploy today, health check failings
a month ago
Update: found that HOSTNAME=0.0.0.0 was set as an explicit Railway service variable (in addition to the Dockerfile's own ENV), which overrides the Dockerfile default at runtime. Tried changing it to HOSTNAME=:: directly as a Railway variable (not just in the Dockerfile) and let the resulting auto-redeploy run — failed identically (14/14 healthcheck attempts, same total silence). Reverted it back to 0.0.0.0 since it didn't help. So that's 6 consecutive failed deployments today across every variant tried (same image redeploy, fresh build, Dockerfile HOSTNAME=::, service-variable HOSTNAME=::). Still looks platform-side to me.
a month ago
Decisive evidence: SSH'd directly into the container while it was mid-healthcheck (railway ssh -s web) and read /proc/net/tcp from inside, since curl/wget aren't in the node:20-slim image:
0: 00000000:0BB8 00000000:0000 0A ...
0BB8 hex = 3000, state 0A = LISTEN, local address 00000000 = bound on all interfaces (0.0.0.0). The app is genuinely listening on the right port, inside the container, at the OS socket level — this isn't a zombie/crashed process despite the silent logs. Whatever's failing is entirely between Railway's healthcheck prober and this listening socket, not in the app.
@ddbruce good to know it's not just us — what service/region are you on? Might help Railway spot the common thread.
a month ago
Found it, and it's on us. On Aug 28 we shipped a security fix that stopped the deploy healthcheck from following redirects. Your healthcheck path / answers with a 307 to /ru, so every attempt now counts as a failure, where before the prober followed it to a 200 and passed, which is why the deployment from two days earlier was fine. Two notes on your evidence: healthcheck probes never appear in the HTTP logs tab since they don't pass through the edge, and the "service unavailable" wording is our misleading hardcoded log message, with a fix in review to report the real status code. Your socket-level check was right, the app was listening fine the whole time. To turn the healthcheck back on, point it at a route that returns 200 directly without the locale redirect, like a dedicated /healthz.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago