a month ago
We have a reproducible Railway deployment-healthcheck failure on a production Next.js frontend.
Service: tgthr-frontend
Region: europe-west4-drams3a
Runtime: Railway Runtime V2
Replicas: 1
Healthcheck: /
Effective app port: 8080
Last successful deployment
Deployment: 987dc2ce-17ee-4e05-b266-dc92a946a1b9
Git: 3f97224311ea62b7e61769f2c32287625e7d5e79
Next startup:
Next.js 15.5.14
Network: http://0.0.0.0:8080
Ready in 572ms
Railway:
Path: /
[1/1] Healthcheck succeeded!
First failing deployment
Deployment: 64e19ba4-c4de-4161-9db7-be7ca85e1a6b
Git: 4ca6b2d3fa530e9c9d3b2189ffe82ed484faa594
Next startup:
Next.js 15.5.14
Network: http://0.0.0.0:8080
Ready in 538ms
Railway then performs 14 healthcheck attempts over ~5 minutes and every attempt returns:
failed with service unavailable
followed by:
1/1 replicas never became healthy!
The process does not crash and remains alive until Railway terminates the failed deployment.
Important A/B comparison
We performed a forensic comparison of LAST GOOD vs FIRST BAD.
They use the same:
Nixpacks 1.41.0
Next.js 15.5.14
Node 22
pnpm 10.13.1
Runtime V2
region
replica count
start command
healthcheck path
0.0.0.0:8080 binding
effective PORT
builder/base image
Railway service manifest
There was no change to package.json, lockfile, Railway config, Nixpacks config, Next config, middleware, Dockerfile, startup scripts, or root routing between the successful deployment and the first failed deployment.
Multiple later deployments now fail identically while the last-good deployment remains active.
We have also ruled out:
missing explicit PORT
target-port mismatch
localhost/bind-host issue
local development port changes
Next/Node/pnpm/Nixpacks version change
image size/resource pressure
/ locale redirect, because the last-good deployment has the same / → 307 /en behavior and passed its healthcheck.
Request
Could Railway please inspect the internal deployment-healthcheck path for these two deployments and compare:
target replica/instance and internal IP
effective healthcheck target port
TCP/connect/upstream-resolution result
HTTP result if the request reached the container
replica/private-network attachment timing/state
compute host/node assignment
routing/upstream errors
Another current Railway report appears to show a very similar issue on Runtime V2 in europe-west4-drams3a, with a healthy application listening on 8080 while deployment healthchecks return service unavailable.
We have intentionally made no application or Railway configuration changes because the previous deployment with the same runtime contract passed successfully.
Screenshots/logs attached.
5 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
This is about as thorough an A/B comparison as I've seen posted here — genuinely nothing obviously wrong on the app side given what you've ruled out. A couple of notes:
This looks like the same underlying pattern as at least one other current thread — "Healthcheck returns 'service unavailable' against a container that is up and idle" describes the identical symptom: process alive, listening correctly, healthcheck still reports unavailable. Worth linking your thread to that one explicitly (and to the "current Railway report" you mention) when you flag this to Railway, since three independent reports of the same Runtime V2/EU-West-region signature is much stronger evidence of a platform-side regression than one report alone — might help get it escalated faster.
One gap not addressed in your diff: you've confirmed no changes across package.json, lockfile, Railway config, Nixpacks config, Next config, middleware, Dockerfile, and startup scripts — but is there a code diff outside those files between the last-good commit and the first-bad commit? Even a change unrelated to routing (e.g. a new import somewhere, an env var read at module scope, a new dependency transitively pulled in without a lockfile version bump if you're not using --frozen-lockfile) could in principle add startup latency or an async operation that shifts when the server is actually ready to accept the healthcheck's specific request, even if Ready in Xms looks nearly identical in the logs. Worth pasting git diff --stat to fully rule that out, since right now the comparison is "these specific config files are unchanged" rather than "nothing changed."
If that diff really is empty too, then this is unambiguously a Runtime V2 platform-side issue and needs Railway's internal healthcheck-path telemetry — which only they can pull.
a month ago
Thank you for your answer. We completed the requested full Git A/B. LAST GOOD 3f972243 → FIRST BAD 4ca6b2d3 is one commit: 57 files, +1159/-1029. Every changed production file reachable from / was inspected. No startup/module-init/network/config/dependency change exists; Next route manifests are semantically identical and both builds generate 108/108 pages. Both candidates start on 0.0.0.0:8080 in ~0.5s.
We also found the independent Runtime V2/EU-West report where explicit PORT=8080 was tested and the same service unavailable persisted, so we are not changing PORT.
Our next single-variable test is changing only healthcheck / (307 → /en) to /en (direct 200). If that still produces service unavailable, we will have eliminated response/redirect semantics too. Can Railway/community confirm whether there is any known Runtime V2 healthcheck/new-replica routing issue in europe-west4-drams3a?
a month ago
UPDATE — root cause isolated: changing only healthcheck / → /en immediately fixed deployment
I changed one setting only:
Healthcheck Path
before: /
after: /en
The next deployment passed its Railway healthcheck successfully.
Nothing else changed:
same Next.js 15.5.14
same Runtime V2
same europe-west4-drams3a
same Nixpacks build
same start command
same 0.0.0.0:8080
same target port
no explicit PORT added
no application-code fix
no Docker change
The relevant application behavior is:
GET / → 307 Location: /en
GET /en → 200
Previous deployments had successfully passed Railway's healthcheck with the same / → 307 /en behavior, including the last-good deployment only ~5 hours before the first failure.
After that point, four consecutive new deployments failed:
Attempt #N failed with service unavailable
despite Next being ready on 0.0.0.0:8080 in ~0.5 seconds.
Changing the healthcheck to the direct-200 /en route immediately resolved the problem.
So this now appears specifically related to Railway healthcheck handling of the redirecting / endpoint rather than the application port, Docker image, Runtime V2 networking, or repository changes.
Railway's current documentation says a healthcheck endpoint must return HTTP 200, which /en does directly while / returns 307.
Open question for Railway: did healthchecker redirect-following/response handling change around Aug 28?The confusing part is that / had passed repeatedly with exactly this 307 redirect before that time.
For now I will keep the direct-200 path and will replace /en with a dedicated lightweight /healthz endpoint so the deployment healthcheck does not depend on a user-facing locale page.
This also disproves our earlier suspected causes:
missing explicit PORT
target-port mismatch
bind host
local dev port changes
Next/Node/Nixpacks changes
image size
application startup failure
If Railway can confirm whether healthcheck redirect handling changed, that would explain the exact last-good → first-bad transition.
a month ago
Direct answer to your open question: yes, healthcheck redirect handling changed on Aug 28. We shipped a security fix that stopped the prober from following redirects, so your / returning a 307 to /en now fails where it used to be followed to the 200 and pass, which is exactly the last-good to first-bad transition you measured. Your fix is right, and the dedicated /healthz you're planning is the ideal target. The "service unavailable" wording is our misleading hardcoded log message, and a fix is in review to report the real status code. We're evaluating whether redirects can be safely supported again, but don't rely on that.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
thanks everyone for support on that topic. Wish you all well in your endeavors!
Status changed to Awaiting Railway Response Railway • about 1 month ago
Status changed to Solved borrkoli2 • about 1 month ago
