a month ago
Subject: Service healthcheck fails ("Application failed to respond" / 502) despite clean, healthy app startup — likely edge↔container networking fault
Project ID: 4c74a3bc-7a0c-4bb8-93ec-85e1032cd6a8 ("passionate-solace")
Environment ID: eda2fb1e-d671-483b-be60-44ee7e667d91 (production)
Service ID: 4db83bb3-e7bc-4bc2-9669-283bf7498514 (name: "fixdiwa")
Public domain: fixdiwa-production.up.railway.app
Recent failing deployment IDs: 9a242c69-e6fc-45c4-8721-dd6dd3151810, ccac705b-586c-4c07-ad43-b623105d762a, 515e8f00-a27e-4027-9223-5039ebba3cec, cd19c68c-63b3-4cf2-9df1-474a82257a2d
Request ID from the "Application failed to respond" page: R3VFN0NlS5W2qCOf9I3ezw
Summary: This service was running fine for a full day (deployment 6a199c77-3be6-4e77-bd87-227f6840a413, commit f9edc498, SUCCESS at 2026-09-01 18:58 UTC, serving real traffic for ~22 hours). Starting with a push at 2026-09-02 17:01 UTC, every deployment since has failed Railway's own healthcheck ("Path: /health", 5-minute retry window, every attempt returns "service unavailable"), and the service now returns "Application failed to respond" / 502 Bad Gateway on the public domain.
What I've already ruled out:
- Application code is not the cause — I redeployed the exact same commit (f9edc498) that ran successfully for a full day, and it failed the identical healthcheck pattern.
- Container startup is clean every single time — deploy logs show migrations complete, "Database connection verified", "Redis rate limiter connected", "API startup complete", "Application startup complete" within ~5 seconds of container start, with zero application errors, on every failed attempt.
- Zero HTTP-layer log entries are ever recorded during any healthcheck attempt (checked via the http log stream on every failed deployment) — the healthcheck probe's requests never appear to reach the application process at all.
- Resource metrics (CPU/memory) during these windows are completely normal — no signs of OOM or resource exhaustion.
- I deleted and recreated the service's public domain (same hostname, fixdiwa-production.up.railway.app) hoping to fix a possibly-stale TLS/routing state — this did not resolve it; the new deployment against the fresh domain failed identically.
- Hitting the raw Railway domain directly returns a clean 502 Bad Gateway (not a TLS/connection error) — suggesting your edge proxy is reachable and can terminate TLS, but cannot successfully reach/relay to the container.
Given the application boots and logs a fully healthy startup every time, but nothing ever reaches it externally, this looks like a fault in the edge↔container networking path (routing table, internal service mesh entry, or similar) specific to this service/deployment, rather than anything in our application. Could you please investigate on your end? Happy to provide any additional deploy/build/http logs you need.
2 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
The public domain you provided is currently working and accessible. Is this issue still happening, or does it happen intermittently?
a month ago
Thanks for checking! Yes, the public domain is working now — but only because we worked around it, not because the underlying issue resolved itself.
What we found: this service's Dockerfile had its own HEALTHCHECK instruction (an in-container loopback probe to 127.0.0.1:8000/health). Clearing Railway's dashboard/API-level healthcheckPath setting had zero effect on deploy behavior across multiple attempts — the logs kept showing "Path: /health" and failing regardless. Removing the HEALTHCHECK instruction from the Dockerfile entirely is what finally let a deploy succeed.
This is a workaround, not a real fix — we still don't know WHY the healthcheck (which had been passing fine for a full day, serving real traffic) suddenly started failing on every single attempt afterward, with the container logging a completely clean, healthy startup every time and zero HTTP-layer requests ever recorded reaching the app during any of the ~11-attempt/5-minute retry windows we saw across 5+ failed deploys.
Given it's resolved for now (workaround in place), this isn't urgent anymore — but we'd genuinely like to understand the root cause before we re-add the HEALTHCHECK instruction, in case it points to something worth knowing about. Happy to share more logs/deployment IDs if useful.