a month ago
Service web (project accomplished-adventure, environment production, domain sviapz-smartcore.up.railway.app). Since 2026-08-28 ~20:24 UTC, every deployment fails its healthcheck: all attempts return "service unavailable" for the whole 5-minute retry window, then "1/1 replicas never became healthy". Four consecutive failed deployments: three of commit 4ba1a30 (Aug 28 20:24, 20:38 UTC; Aug 29 morning) and one of commit e938816 (Aug 29 ~08:45 UTC) — two different builds, same result.
In every attempt the Deploy Logs show the container completing startup normally: migrations, collectstatic, then gunicorn 25.1.0 logging Listening at: http://0.0.0.0:8080 and booting its worker — yet healthcheck attempts keep failing after the app is listening. Config-as-code (railway.toml) is unchanged since its introduction: healthcheckPath = "/", healthcheckTimeout = 300; target port 8080; no PORT/variable changes; no incidents on the status page. The last successful deployment was commit f984131 on 2026-08-28 ~07:08 UTC with the exact same configuration — the currently active deployment, still serving traffic correctly, which also shows the app itself responds fine once routed.
Could you check whether something changed on the internal healthcheck path/networking for this service after Aug 28 ~07:00 UTC? Deployment IDs available on request from the dashboard history.
6 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
This matches the exact signature of at least two other current threads — "Runtime V2 EU West: new frontend deploys fail healthcheck while previous identical deployment stays healthy" (borrkoli2) and "Healthcheck returns 'service unavailable' against a container that is up and idle" — all three describe: app logs confirm the process is listening correctly, config is unchanged from a previously-working deployment, yet the healthcheck reports unavailable for the full retry window. Worth explicitly cross-linking your thread to those when you flag this, since the pattern starting specifically around Aug 28 evening onward across multiple independent services is a much stronger signal of a Runtime V2 platform-side regression than any single report — that's worth stating clearly to Railway.
One thing not yet mentioned in your report that's worth adding: what region is this service in? The other two threads are both flagged as EU West (europe-west4-drams3a / drams3a-adjacent). If yours is also EU West (or also happens to be US-based, which would broaden the blast radius), that detail alone would help Railway correlate faster — right now your post doesn't state the region explicitly.
Also worth noting for the record: your healthcheck path is / with a 300s timeout, and the last-good deployment used the exact same config — so this isn't a case of a slow app needing a longer timeout, which rules out the most common self-inflicted cause of this error. Combined with the other reports, this looks like it needs Railway to check the internal healthcheck-routing path for Runtime V2 broadly around the Aug 28 ~20:00 UTC mark, not just this one service.
a month ago
Adding details for cross-correlation, as suggested:
• Region: US West — notably NOT EU. Per the point raised above: a US-based case with the same signature broadens the blast radius beyond the EU West cluster.
• Timing: our first two failed deploys were at 20:24 and 20:38 UTC on Aug 28 — within ~30 minutes of the ~20:00 UTC onset reported for the EU West cases. Same signature, same time window, different region: this points at a platform-wide rollout (Runtime V2 / healthcheck delivery path) rather than a region-specific network fault.
• Cross-links: same signature as [link thread borrkoli2] and [link thread "container up and idle"] — process verifiably listening, config unchanged from a previously green deployment (last green: Aug 28 ~07:08 UTC), healthcheck "service unavailable" for the entire retry window.
• Healthcheck path is / with a 300s timeout, identical to the last-good deployment — not a slow-startup/timeout case.
a month ago
Good to add the US-West data point — that does argue for a platform-wide rollout rather than a regional fault, given the ~30min window overlap with the EU cases.
One more thing worth checking: someone with the identical symptom (EU West, Next.js, same "listening but service unavailable" signature) found their fix by testing whether their healthcheck route returns a redirect instead of a direct 200. Their / route was doing a 307 redirect to /en, and switching the healthcheck path to a route that returns 200 directly fixed it immediately, with nothing else changed.
Does your / route also redirect (e.g. to a locale, an auth check, or similar), or does it return 200 directly? If yours also redirects, try pointing the healthcheck at a route that returns 200 directly (or add a lightweight /healthz) and redeploy. If that fixes it too, that's strong confirmation across three independent services/regions that this is specifically about Runtime V2's healthcheck handling of redirect responses — not a regional network issue.
If your / doesn't redirect, that's useful too — it'd argue against redirect-handling being the full explanation.
a month ago
CONFIRMED — third independent service. Our / returns a 302 to /login/ (Django auth). It passed healthchecks for months, last green Aug 28 ~07:08 UTC, then every deploy failed from ~20:24 UTC onward with zero config changes. Just switched healthcheckPath to /login/ (a plain direct 200): attempt #1 failed while the app was still starting, attempt #2 succeeded 7 seconds later. Deployment green, nothing else changed in that commit. That's now three services across US West and EU West (302 auth / 307 locale / 308 markets), same onset window — this is a Runtime V2 prober regression in redirect handling, starting Aug 28 ~20:00 UTC. Railway: redirect-following healthchecks silently stopped working; this needs a platform fix or at minimum a documented breaking change.
a month ago
This closes the loop nicely — three independent services, three different redirect status codes (302/307/308), two regions, all failing in the same ~20-minute window on Aug 28 and all fixed by the same single change (healthcheck path returns 200 directly instead of following/receiving a redirect). That rules out anything code- or region-specific; the common factor is exclusively "healthcheck path redirects" + "deployed/redeployed after ~Aug 28 20:00 UTC."
Summarizing for anyone landing on this thread cluster: if your deployment health-checked fine before Aug 28 evening and started failing since with no config changes, and your healthcheck path returns a redirect (any 3xx) rather than a direct 200 — point the healthcheck at a route that returns 200 directly (or add a dedicated /healthz) and redeploy. That's the confirmed workaround across at least three unrelated services now.
If this is what fixed it for you, could you mark this reply as the accepted solution? Thanks
a month ago
Confirming your diagnosis: on Aug 28 we shipped a security fix that stopped the deploy healthcheck from following redirects, in exactly the window you pinned down. A 3xx from the healthcheck path now counts as a failure, where before the prober followed it to the 200. Your fix is the right one, and a dedicated lightweight endpoint is the ideal long-term target. On our side, the "service unavailable" wording is a hardcoded log message, and a fix is in review to report the real status code, which would have pointed you at the redirect immediately. We're also evaluating whether redirects can be safely supported again, but don't rely on that.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago
