a month ago
Every new deployment triggered by a push to main fails during the Network Healthcheck step, always at almost exactly 95 seconds, well under the 120s healthcheck timeout in railway.json.
The app itself starts correctly and quickly, listening on 0.0.0.0:8080 within about 2 seconds per deploy logs. But no HTTP request ever reaches the app during the failure window. We log every incoming request with Serilog request logging, and there are zero entries for the healthcheck path during the roughly 95 seconds the probe is supposedly running. The Network Logs tab in the dashboard shows no HTTP entries either.
This has happened on essentially every deploy for the last 24+ hours. Only one deployment succeeded in that window and is still ACTIVE serving traffic. I manually redeployed the same failing commit twice and both attempts failed at the exact same 95 second mark in the healthcheck step.
Ruled out on my end: port config matches (Networking target port 8080, app binds to 8080 via PORT fallback), the health path is exempt from our rate limiter, resource limits are far from the cap (8 vCPU 8GB allocated, usage stays low), Outbound IPv6 is disabled, and Private Networking shows Ready to talk privately. I also tried the built in Diagnose AI tool twice on a failed deployment and it errored out both times without producing a result.
Project: heroic-spontaneity, Service: customerAI, Environment: production, Region: EU West europe-west4-drams3a. Failed deployment id: 0427ee89-ebf4-497b-8feb-313f1667555b, plus the redeploy right after it on the same commit.
I noticed other threads on this region reporting similar symptoms, e.g. one 7 months ago titled europe-west4-drams3a instance not receiving any HTTP request, and a Railway employee confirmed a separate ongoing incident on this region a month ago (deploys blocked region service unreachable). Wondering if this is another instance of region level networking trouble. Would appreciate help figuring out why the healthcheck probe is not reaching the container network, since the app is verifiably healthy.
2 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
This does not look like an application-level healthcheck failure. The strongest evidence is that the container is already listening on 0.0.0.0:8080, yet no healthcheck request reaches the application or appears in Network Logs during the failure window.
Railway healthchecks are supposed to query the configured endpoint until they receive HTTP 200, using the service PORT. If your railway.json timeout is 120s, consistently failing at ~95s also suggests the failure is happening before the application-level timeout path completes.
There are prior reports with nearly the same signature in europe-west4-drams3a: healthy container, same commit intermittently succeeding, healthcheck traffic never reaching the container, and failures isolated to Railway routing/service registration.
I would treat this as a regional deployment-network / service-registration issue rather than changing the application.
Two useful confirmation tests:
Temporarily deploy the same service/commit in another Railway region. If it passes there without code/config changes, that strongly isolates the fault to europe-west4-drams3a.
On a failing deployment, compare the configured 120s timeout with the actual ~95s failure time. If Railway aborts consistently before the configured timeout while no packet reaches the container, include that plus the deployment IDs in a Railway support escalation.
I would not change the bind address, health path, resource limits, or application code based on the evidence shown; those do not explain why the probe never reaches the container.
Railway staff will likely need to inspect the internal healthcheck routing / service registration for deployment 0427ee89-ebf4-497b-8feb-313f1667555b and the immediately following redeploy in europe-west4-drams3a.
a month ago
Not sure this is the exact same issue you're hitting, but it read close enough to what we just went through that I figured I'd mention it in case it helps someone landing here.
We also had every deploy die at the healthcheck step with a SIGTERM, "service unavailable" x14. But in our case the logs showed the app actually reaching "Ready" and the healthcheck request WAS reaching the container — our healthcheck path (/) was just behind our own auth middleware, which redirects any unauthenticated request to the login page (307), so Railway's probe never got the 200 it was waiting for. So for us it turned out to be entirely application-side — we'd pointed the healthcheck at an auth-gated route by mistake.
What you're describing (zero entries reaching the container, nothing in Network Logs during the failure window) sounds like a genuinely different signature — much more consistent with a region/networking issue like turicamirabelamaria-art suggested. Our fix (pointing the healthcheck at a dedicated unauthenticated endpoint) almost certainly wouldn't help your case since the request isn't reaching the app at all in your logs. Just thought the surface-level similarity was worth flagging in case it's useful context.
Hope it gets sorted quickly — ours at least turned out to be an obvious app-side mistake once we found it; yours looks like the more frustrating kind.