a month ago
Project: savaria-research (d3fc7e3c-8112-4b65-b965-31af3bea92fd) · Service: savaria-research (6450b2fa-4955-49b6-b9c4-54ccea0a2814) · Environment: production (112e67ca-18a8-44bf-8c7e-09959e1189f7)
Summary:
Every deploy on this service has failed since 2026-08-28 19:37 (CEST) — over a dozen attempts, roughly 18 hours. Symptom is always identical: the Next.js app boots and logs "Ready" within ~300-600ms, then the container produces zero further log output until Railway kills it once the healthcheck window (tested at both 30s and 180s) expires. Healthcheck path is /, confirmed matching what's configured in railway.json via the dashboard.
Proof this is not our application:
Redeployed the exact commit that succeeded on 2026-08-28 (3a38304) with zero code changes, via a throwaway branch. It failed identically.
Added instrumentation directly inside the running Node process — a live census of open sockets/timers and active HTTP requests, logged every 5 seconds from boot. Result: 0 active requests, logged identically every 5 seconds across the entire 180-second healthcheck window, right up until the container was killed. The process was fully healthy, idle, and correctly listening on its port the whole time — it never received the healthcheck's HTTP request at all.
Already ruled out, with direct evidence, not assumption:
Memory — usage never exceeded ~700MB even under the original 1GB limit (checked via railway metrics --memory --raw)
DB reachability — fast, healthy responses to our database from both our own machines and from inside Railway's network via railway ssh
DNS/IPv6 — no AAAA record exists for our DB host, so an IPv6-hang class of bug is structurally impossible here
Postgres-side locking — checked live via pg_stat_activity during an active failure window, nothing non-idle
Healthcheck config drift — dashboard confirmed path/timeout match railway.json exactly
Disabling the healthcheck entirely (done via your own Railway AI agent) — still failed the same way, container SIGTERM'd before completing its own startup, proving the healthcheck itself was never the actual gate
A local production build against our real production database works correctly every time
Ask: Given the app is proven healthy and idle, and the healthcheck request itself never arrives, this looks like something in the request-delivery path specific to this service (edge routing, internal DNS, proxy state) rather than anything in our container. Could someone from Railway look at this service's deploy pipeline directly? The site itself (research.savaria.in) has stayed up throughout, served by the last successful deployment — this is blocking new deploys, not causing a live outage, but has now blocked us ~18 hours.
Pinned Solution
a month ago
Root cause
The service's healthcheckPath was set to /, which returns a 308 redirect (to /markets/in) rather than serving content directly. Railway's healthcheck prober did not handle this redirect correctly for this service — it never delivered a request that the application could respond to, so every deploy's healthcheck failed and the container was killed once the retry window elapsed, regardless of any code change.
Fix: repointed healthcheckPath to a dedicated endpoint that returns a plain, direct 200 with no redirect involved. Deploys have succeeded reliably since.
How this was isolated
The investigation initially assumed the cause was application-level (a database statement timeout was genuinely found and fixed early on, but turned out to be unrelated to the deploy failures). Each hypothesis below was tested directly and ruled out with evidence, not assumption, before the actual cause was found:
Memory — railway metrics --memory --raw showed real usage never exceeded ~700MB, even under the original 1GB limit. No OOM.
Application code — redeployed the exact commit that had succeeded 24 hours earlier, with zero changes. It failed identically, ruling out every commit in the repository.
Database reachability — confirmed fast and healthy from both outside Railway and from inside Railway's own network (via railway ssh).
DNS/IPv6 — no AAAA record exists for the database host, ruling out an IPv6-hang class of bug.
Postgres-side locking — checked live via pg_stat_activity during an active failure window; nothing non-idle.
Deploy trigger mechanism — GitHub-webhook-triggered deploys, railway redeploy, and railway up all failed identically (the last ran 8+ minutes before being killed, well past the healthcheck window).
Custom domain — temporarily detached, tested with only the Railway-generated domain. Same failure.
Account/workspace-level fault — ruled out because a sibling service in the same project/workspace/account deployed successfully the entire time.
The healthcheck itself being the blocker — disabling it entirely (via Railway's own AI agent) still resulted in the container being killed shortly after starting, which was the clue that misdirected the investigation briefly before the redirect-handling theory was tested directly.
The decisive diagnostic: instrumentation added directly inside the running Node process (a live census of open sockets and active HTTP requests, logged every 5 seconds) showed zero incoming requests for the entire healthcheck window, proving the app was healthy and idle the whole time — the request simply never arrived. Comparing against the sibling service (which healthchecks a plain, non-redirecting endpoint and always worked) pointed directly at the redirect as the remaining variable.
Resolution confirmed
Live verification post-fix: / returns 200 (0.73s), /markets/in returns 200 directly (0.5s), and the new healthcheck endpoint returns 200 (0.26s) — all on the production custom domain.
3 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
Same issue with me. Doesn't even redeploy successfully what was already deployed.
a month ago
Root cause
The service's healthcheckPath was set to /, which returns a 308 redirect (to /markets/in) rather than serving content directly. Railway's healthcheck prober did not handle this redirect correctly for this service — it never delivered a request that the application could respond to, so every deploy's healthcheck failed and the container was killed once the retry window elapsed, regardless of any code change.
Fix: repointed healthcheckPath to a dedicated endpoint that returns a plain, direct 200 with no redirect involved. Deploys have succeeded reliably since.
How this was isolated
The investigation initially assumed the cause was application-level (a database statement timeout was genuinely found and fixed early on, but turned out to be unrelated to the deploy failures). Each hypothesis below was tested directly and ruled out with evidence, not assumption, before the actual cause was found:
Memory — railway metrics --memory --raw showed real usage never exceeded ~700MB, even under the original 1GB limit. No OOM.
Application code — redeployed the exact commit that had succeeded 24 hours earlier, with zero changes. It failed identically, ruling out every commit in the repository.
Database reachability — confirmed fast and healthy from both outside Railway and from inside Railway's own network (via railway ssh).
DNS/IPv6 — no AAAA record exists for the database host, ruling out an IPv6-hang class of bug.
Postgres-side locking — checked live via pg_stat_activity during an active failure window; nothing non-idle.
Deploy trigger mechanism — GitHub-webhook-triggered deploys, railway redeploy, and railway up all failed identically (the last ran 8+ minutes before being killed, well past the healthcheck window).
Custom domain — temporarily detached, tested with only the Railway-generated domain. Same failure.
Account/workspace-level fault — ruled out because a sibling service in the same project/workspace/account deployed successfully the entire time.
The healthcheck itself being the blocker — disabling it entirely (via Railway's own AI agent) still resulted in the container being killed shortly after starting, which was the clue that misdirected the investigation briefly before the redirect-handling theory was tested directly.
The decisive diagnostic: instrumentation added directly inside the running Node process (a live census of open sockets and active HTTP requests, logged every 5 seconds) showed zero incoming requests for the entire healthcheck window, proving the app was healthy and idle the whole time — the request simply never arrived. Comparing against the sibling service (which healthchecks a plain, non-redirecting endpoint and always worked) pointed directly at the redirect as the remaining variable.
Resolution confirmed
Live verification post-fix: / returns 200 (0.73s), /markets/in returns 200 directly (0.5s), and the new healthcheck endpoint returns 200 (0.26s) — all on the production custom domain.
Root cause The service's healthcheckPath was set to /, which returns a 308 redirect (to /markets/in) rather than serving content directly. Railway's healthcheck prober did not handle this redirect correctly for this service — it never delivered a request that the application could respond to, so every deploy's healthcheck failed and the container was killed once the retry window elapsed, regardless of any code change. Fix: repointed healthcheckPath to a dedicated endpoint that returns a plain, direct 200 with no redirect involved. Deploys have succeeded reliably since. How this was isolated The investigation initially assumed the cause was application-level (a database statement timeout was genuinely found and fixed early on, but turned out to be unrelated to the deploy failures). Each hypothesis below was tested directly and ruled out with evidence, not assumption, before the actual cause was found: Memory — railway metrics --memory --raw showed real usage never exceeded ~700MB, even under the original 1GB limit. No OOM. Application code — redeployed the exact commit that had succeeded 24 hours earlier, with zero changes. It failed identically, ruling out every commit in the repository. Database reachability — confirmed fast and healthy from both outside Railway and from inside Railway's own network (via railway ssh). DNS/IPv6 — no AAAA record exists for the database host, ruling out an IPv6-hang class of bug. Postgres-side locking — checked live via pg_stat_activity during an active failure window; nothing non-idle. Deploy trigger mechanism — GitHub-webhook-triggered deploys, railway redeploy, and railway up all failed identically (the last ran 8+ minutes before being killed, well past the healthcheck window). Custom domain — temporarily detached, tested with only the Railway-generated domain. Same failure. Account/workspace-level fault — ruled out because a sibling service in the same project/workspace/account deployed successfully the entire time. The healthcheck itself being the blocker — disabling it entirely (via Railway's own AI agent) still resulted in the container being killed shortly after starting, which was the clue that misdirected the investigation briefly before the redirect-handling theory was tested directly. The decisive diagnostic: instrumentation added directly inside the running Node process (a live census of open sockets and active HTTP requests, logged every 5 seconds) showed zero incoming requests for the entire healthcheck window, proving the app was healthy and idle the whole time — the request simply never arrived. Comparing against the sibling service (which healthchecks a plain, non-redirecting endpoint and always worked) pointed directly at the redirect as the remaining variable. Resolution confirmed Live verification post-fix: / returns 200 (0.73s), /markets/in returns 200 directly (0.5s), and the new healthcheck endpoint returns 200 (0.26s) — all on the production custom domain.
a month ago
Thanks, this has worked for my case! build is green! Thanks once again! Cheers.
Status changed to Solved medim • about 1 month ago
