Runtime V2 EU West: new frontend deploys fail healthcheck while previous identical deployment stays healthy
borrkoli2
HOBBYOP

a month ago

We have a reproducible Railway deployment-healthcheck failure on a production Next.js frontend.

Service: tgthr-frontend

Region: europe-west4-drams3a

Runtime: Railway Runtime V2

Replicas: 1

Healthcheck: /

Effective app port: 8080

Last successful deployment

Deployment: 987dc2ce-17ee-4e05-b266-dc92a946a1b9

Git: 3f97224311ea62b7e61769f2c32287625e7d5e79

Next startup:

Next.js 15.5.14

Network: http://0.0.0.0:8080

Ready in 572ms

Railway:

Path: /

[1/1] Healthcheck succeeded!

First failing deployment

Deployment: 64e19ba4-c4de-4161-9db7-be7ca85e1a6b

Git: 4ca6b2d3fa530e9c9d3b2189ffe82ed484faa594

Next startup:

Next.js 15.5.14

Network: http://0.0.0.0:8080

Ready in 538ms

Railway then performs 14 healthcheck attempts over ~5 minutes and every attempt returns:

failed with service unavailable

followed by:

1/1 replicas never became healthy!

The process does not crash and remains alive until Railway terminates the failed deployment.

Important A/B comparison

We performed a forensic comparison of LAST GOOD vs FIRST BAD.

They use the same:

Nixpacks 1.41.0

Next.js 15.5.14

Node 22

pnpm 10.13.1

Runtime V2

region

replica count

start command

healthcheck path

0.0.0.0:8080 binding

effective PORT

builder/base image

Railway service manifest

There was no change to package.json, lockfile, Railway config, Nixpacks config, Next config, middleware, Dockerfile, startup scripts, or root routing between the successful deployment and the first failed deployment.

Multiple later deployments now fail identically while the last-good deployment remains active.

We have also ruled out:

missing explicit PORT

target-port mismatch

localhost/bind-host issue

local development port changes

Next/Node/pnpm/Nixpacks version change

image size/resource pressure

/ locale redirect, because the last-good deployment has the same / → 307 /en behavior and passed its healthcheck.

Request

Could Railway please inspect the internal deployment-healthcheck path for these two deployments and compare:

target replica/instance and internal IP

effective healthcheck target port

TCP/connect/upstream-resolution result

HTTP result if the request reached the container

replica/private-network attachment timing/state

compute host/node assignment

routing/upstream errors

Another current Railway report appears to show a very similar issue on Runtime V2 in europe-west4-drams3a, with a healthy application listening on 8080 while deployment healthchecks return service unavailable.

We have intentionally made no application or Railway configuration changes because the previous deployment with the same runtime contract passed successfully.

image.png

image.png

Screenshots/logs attached.

Solved$10 Bounty

5 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


dolapobayo42-lgtm
FREE

a month ago

This is about as thorough an A/B comparison as I've seen posted here — genuinely nothing obviously wrong on the app side given what you've ruled out. A couple of notes:

This looks like the same underlying pattern as at least one other current thread — "Healthcheck returns 'service unavailable' against a container that is up and idle" describes the identical symptom: process alive, listening correctly, healthcheck still reports unavailable. Worth linking your thread to that one explicitly (and to the "current Railway report" you mention) when you flag this to Railway, since three independent reports of the same Runtime V2/EU-West-region signature is much stronger evidence of a platform-side regression than one report alone — might help get it escalated faster.

One gap not addressed in your diff: you've confirmed no changes across package.json, lockfile, Railway config, Nixpacks config, Next config, middleware, Dockerfile, and startup scripts — but is there a code diff outside those files between the last-good commit and the first-bad commit? Even a change unrelated to routing (e.g. a new import somewhere, an env var read at module scope, a new dependency transitively pulled in without a lockfile version bump if you're not using --frozen-lockfile) could in principle add startup latency or an async operation that shifts when the server is actually ready to accept the healthcheck's specific request, even if Ready in Xms looks nearly identical in the logs. Worth pasting git diff --stat to fully rule that out, since right now the comparison is "these specific config files are unchanged" rather than "nothing changed."

If that diff really is empty too, then this is unambiguously a Runtime V2 platform-side issue and needs Railway's internal healthcheck-path telemetry — which only they can pull.


borrkoli2
HOBBYOP

a month ago

Thank you for your answer. We completed the requested full Git A/B. LAST GOOD 3f972243 → FIRST BAD 4ca6b2d3 is one commit: 57 files, +1159/-1029. Every changed production file reachable from / was inspected. No startup/module-init/network/config/dependency change exists; Next route manifests are semantically identical and both builds generate 108/108 pages. Both candidates start on 0.0.0.0:8080 in ~0.5s.

We also found the independent Runtime V2/EU-West report where explicit PORT=8080 was tested and the same service unavailable persisted, so we are not changing PORT.

Our next single-variable test is changing only healthcheck / (307 → /en) to /en (direct 200). If that still produces service unavailable, we will have eliminated response/redirect semantics too. Can Railway/community confirm whether there is any known Runtime V2 healthcheck/new-replica routing issue in europe-west4-drams3a?


borrkoli2
HOBBYOP

a month ago

UPDATE — root cause isolated: changing only healthcheck / → /en immediately fixed deployment

I changed one setting only:

Healthcheck Path

before: /

after: /en

The next deployment passed its Railway healthcheck successfully.

Nothing else changed:

same Next.js 15.5.14

same Runtime V2

same europe-west4-drams3a

same Nixpacks build

same start command

same 0.0.0.0:8080

same target port

no explicit PORT added

no application-code fix

no Docker change

The relevant application behavior is:

GET / → 307 Location: /en

GET /en → 200

Previous deployments had successfully passed Railway's healthcheck with the same / → 307 /en behavior, including the last-good deployment only ~5 hours before the first failure.

After that point, four consecutive new deployments failed:

Attempt #N failed with service unavailable

despite Next being ready on 0.0.0.0:8080 in ~0.5 seconds.

Changing the healthcheck to the direct-200 /en route immediately resolved the problem.

So this now appears specifically related to Railway healthcheck handling of the redirecting / endpoint rather than the application port, Docker image, Runtime V2 networking, or repository changes.

Railway's current documentation says a healthcheck endpoint must return HTTP 200, which /en does directly while / returns 307.

Open question for Railway: did healthchecker redirect-following/response handling change around Aug 28?The confusing part is that / had passed repeatedly with exactly this 307 redirect before that time.

For now I will keep the direct-200 path and will replace /en with a dedicated lightweight /healthz endpoint so the deployment healthcheck does not depend on a user-facing locale page.

This also disproves our earlier suspected causes:

missing explicit PORT

target-port mismatch

bind host

local dev port changes

Next/Node/Nixpacks changes

image size

application startup failure

If Railway can confirm whether healthcheck redirect handling changed, that would explain the exact last-good → first-bad transition.


sam-a
EMPLOYEE

a month ago

Direct answer to your open question: yes, healthcheck redirect handling changed on Aug 28. We shipped a security fix that stopped the prober from following redirects, so your / returning a 307 to /en now fails where it used to be followed to the 200 and pass, which is exactly the last-good to first-bad transition you measured. Your fix is right, and the dedicated /healthz you're planning is the ideal target. The "service unavailable" wording is our misleading hardcoded log message, and a fix is in review to report the real status code. We're evaluating whether redirects can be safely supported again, but don't rely on that.


Status changed to Awaiting User Response Railway • about 1 month ago


borrkoli2
HOBBYOP

a month ago

thanks everyone for support on that topic. Wish you all well in your endeavors!


Status changed to Awaiting Railway Response Railway • about 1 month ago


Status changed to Solved borrkoli2 • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...