Healthcheck returns "service unavailable" against a container that is up and idle.
alexpopescufl
PROOP

a month ago

Healthcheck returns "service unavailable" against a container that is up and idle.

Project: finstar-stage — 21857105-831c-465f-abc0-9daa0704f257

Service: finstar-stage-frontend — 3140ed0a-9b09-4476-92af-a71cc6131a0d

Environment: stage-env — 74390988-7d37-412f-9ff7-fa753c299daf

Region europe-west4-drams3a, Railpack, runtime V2, 1 replica, Next.js 16.2.6, custom domain stage.finstar.ro to target port 8080.

SYMPTOM

Deploys fail the healthcheck while the container is up and listening on the right port. Every probe returns "failed with service unavailable" and the app logs never record a single request. This has happened ~14 times since 2026-05-27 on unrelated commits, then recovered on its own. On 2026-08-28 it failed five times in a row across three different nodes, while every sibling service deployed fine from the same commits.

Failed deploys 2026-08-28: 9d6ca3c6-8bac-43fc-92c7-98ddc36e039b, d88d0193-a48b-446e-8381-139509382a55, 35c0da1f-13d1-41c0-ac81-adb20ac772c5, 147772f6-18f4-44cb-b2ba-2d04a2034558, 20ee745e-212b-45a4-92af-d65bd884017c.

Successful one for contrast: df0fde53-dc01-4c17-a9cb-ec95e3ff1cdf at 16:05 — passed on attempt #1 in 0.56s.

The whole deploy log is fifteen lines: Starting Container, then "Next.js 16.2.6 / Ready in 239ms" binding 8080, then seven minutes of silence, then Stopping Container and a SIGTERM teardown. The http stream is empty. Memory peaked at 0.26 GB of a 24 GB limit; CPU stayed at zero.

RULED OUT, WITH THE CHANGE THAT RULED IT OUT

  • Our code. The commit that failed touches only client components behind authenticated routes; the root path's server bundle is 141 modules and none changed. Our middleware is synchronous and redirects an anonymous GET / without rendering.
  • A hung process. A hung server still completes the TCP handshake, which would time out rather than report "service unavailable". CPU zero and flat memory rule out a blocked event loop.
  • The window. Raised healthcheckTimeout 30 to 300. The next deploy made thirteen attempts over four minutes, every one refused.
  • IPv4-only bind. Changed to --hostname :: — container log confirms Local http://[::1]:8080 and Network http://[::]:8080. Ten more refusals.
  • The 307. / redirects to /login. You have a distinct "failed with status N" wording and have never used it here; successful deploys pass on the same redirect.
  • Node or platform fault. Three different nodes; status page reports 100% regional uptime across the 90 days containing every failure.

THE DECISIVE EXPERIMENT

We removed the healthcheck. The deploy succeeded and stage.finstar.ro immediately served real traffic from it — GET / returns 307, /login returns 200 in 0.29s. We then restored the healthcheck unchanged against that same running service (deploy a47a4b50-e6e7-463c-be26-f5e8cc9e8035, 19:07). It failed again, ten refusals.

That is what makes this specific. The prober is not interrogating the container that is serving. It interrogates the new replica, which has not been attached to the network yet, while an older attached replica handles public traffic throughout. "The app answers the internet but not your probe" is not a contradiction — those are two different containers at two different stages.

It also explains the three-month intermittency. When attachment completes before the probe, the check passes on attempt #1 in half a second. When it does not, retrying cannot help, because the replica is never reachable at all.

QUESTIONS

  1. What does "failed with service unavailable" mean at your routing layer here — the prober failing to resolve the replica as an upstream, or failing to connect to it?
  2. Is there a known attach delay or fault for new replicas under runtime V2 in europe-west4-drams3a?
  3. Is the healthcheck.railway.app Host header relevant for a Next.js app that does no host filtering?

We have had to disable the healthcheck on this service to unblock deploys and would like to reverse that. Happy to leave it as is for inspection, or re-enable it on request so you can watch a failure live.

Solved$20 Bounty

4 Replies

Railway
BOT

a month ago

We've looked into this from our side and haven't found anything on the Railway platform that explains what you're seeing, so working it out means digging into your specific setup.

That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway • about 1 month ago


alexpopescufl
PROOP

a month ago

Another user reports the same thing today, also on staging: station.railway.com/questions/getting-502-gateway-errors-on-existing-d-09d61173. We changed nothing in our network configuration, the injected PORT resolves to 8080 in the container log, the app binds 8080, the domain targets 8080, and we bound both stacks with --hostname :: and confirmed [::]:8080. It still fails. The workaround suggested there — remove the healthcheck — is what we independently arrived at, and it works, which is itself evidence the app is fine and the probe path is not.

UPDATE: THE PORT ANSWER IS NOW TESTED AND DEAD

Following the advice given on the other thread, we set PORT=8080 explicitly as a service variable, matching the domain's target port, and restored the healthcheck. Deployment ab707132-8ce5-4530-b076-6ba96961aabf, 19:37: app reports "Ready in 106ms" on [::]:8080, and the probe fails nine times with the same "service unavailable". So a missing PORT variable is not the cause either. The variable is left in place; it is harmless and correct.

QUESTIONS

What does "failed with service unavailable" mean at your routing layer here — the prober failing to resolve the replica as an upstream, or failing to connect to it?

Is there a known attach delay or fault for new replicas under runtime V2 in europe-west4-drams3a?

Is the healthcheck.railway.app Host header relevant for a Next.js app that does no host filtering?

We have had to disable the healthcheck on this service to unblock deploys and would like to reverse that. Happy to leave it as is for inspection, or re-enable it on request so you can watch a failure live.


Is the healthcheck path set to a route that returns 200 as the HTML status code (e.g., /health)?


sam-a
EMPLOYEE

a month ago

The 307 you ruled out is the cause, and the clue that made you rule it out is our bug. On Aug 28 we shipped a security fix that stopped the deploy healthcheck from following redirects, so a 3xx from the healthcheck path now fails every attempt where it used to be followed to a 200 and pass. The "failed with status N" wording you expected doesn't exist in this log path, "service unavailable" is hardcoded for every failure, and a fix is in review to report the real status code. That's why deploys with the same redirect passed before Aug 28 and why removing the healthcheck unblocked you. There's no replica attach fault, and the Host header isn't relevant here. To turn the healthcheck back on, point it at /login or a dedicated /healthz that returns 200 directly. The intermittent failures you saw from May onward are a separate story, happy to dig in if they return, but the deterministic failures since Aug 28 are all this.


Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...