Healthcheck reports service unavailable, but no healthcheck connections reach running instance
nirtituani
HOBBYOP

a month ago

Hi Railway team,

I'm experiencing repeated healthcheck failures on a production frontend service.

The container starts successfully and remains RUNNING during the entire healthcheck window. The healthcheck then reports:

Path: /

Attempt #1 ... failed with service unavailable

...

Attempt #14 ... failed with service unavailable

Healthcheck failed!

The interesting part is that I instrumented the running container during the healthcheck window.

While Railway was reporting the healthcheck failures, I observed zero inbound connections to the application's listening port (8080) during the entire observation period.

As a control, I ran the exact same read-only network observation against a previously healthy instance of the same service. It reliably detected incoming Railway proxy connections there.

I also verified on the failing instance that:

  • The container is RUNNING.
  • The application is listening on [::]:8080.
  • localhost/loopback requests return HTTP 307.
  • The application is reachable from another service in the same Railway project and returns HTTP 307.
  • Explicitly setting PORT=8080 did not change the behavior.
  • Setting the domain target port to 8080 did not change the behavior.
  • We also tested different hostname bindings, including 0.0.0.0 and ::, without changing the healthcheck result.

The last known-good deployment was on August 26. Multiple deployments since then have failed in exactly the same way, despite the container starting normally.

At this point, the evidence seems to indicate that the healthcheck request may not be reaching the deployment instance at all.

Could someone help clarify:

  1. Which address/port does Railway's healthcheck target for a service like this?
  2. Is the healthcheck expected to connect directly to the deployment instance?
  3. Is there any known issue with healthcheck routing that could cause service unavailable when the container itself is reachable?

I can provide deployment IDs, timestamps, logs, and additional network evidence privately if needed.

Thanks!

Solved$10 Bounty

2 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


dolapobayo42-lgtm
FREE

a month ago

The packet-level instrumentation here is the most direct evidence in any of these related threads — you're not just inferring "the healthcheck isn't reaching it," you've actually observed zero inbound connections on the listening port during the failure window, with a live control (same service, previously healthy) showing connections arriving normally. That's about as close to proof as you can get from outside the platform that this is a proxy/routing issue, not an app-side one.

Two things worth adding to strengthen this before Railway looks at it:

Cross-link to the other open healthcheck threads — there's a cluster of very similar reports right now: "Runtime V2 EU West: new frontend deploys fail healthcheck..." (borrkoli2) and gmazzone79-ctrl's report (US West), both starting around Aug 28. borrkoli2 isolated their fix to the healthcheck path returning a 307 redirect instead of a direct 200 — you mention your app also returns 307 on / (per your loopback test). Worth explicitly testing: point the healthcheck at a route that returns 200 directly and see if it resolves, same as their fix. If it does, that ties your case to the same underlying bug even though your symptom (zero inbound connections) looks more severe than theirs (connections arrived but got rejected as unhealthy) — possibly a broader/harder failure mode of the same regression.

If the direct-200 test does not fix it, that's actually a useful negative result — it'd mean your case is a distinct routing failure (proxy never dispatching the healthcheck at all) rather than the redirect-handling bug, and worth explicitly separating from the other threads so Railway doesn't conflate two different bugs under one fix.

Given your last-known-good was Aug 26 (slightly earlier than the ~Aug 28 20:00 UTC onset in the other reports) and you have the strongest direct network evidence of the group, I'd suggest sharing the deployment IDs/timestamps/network capture with Railway support directly rather than waiting on community diagnosis — this looks like it needs someone with access to the proxy's actual routing decisions for your instance.


sam-a
EMPLOYEE

a month ago

Answers to your three questions: the healthcheck connects directly to the new deployment's container on its PORT, it never goes through the proxy or your domain, and yes, there's a known change. On Aug 28 we shipped a security fix that stopped the healthcheck from following redirects. Your / returns a 307, which now counts as a failure on every attempt, where before the prober followed it to a 200 and passed. That also reframes your control test: the connections you saw on the healthy instance were live proxy traffic, which the healthcheck never uses, and the probe's connections are brief enough that polling-based observation can miss them. The "service unavailable" wording is our misleading hardcoded log message, and a fix is in review to show the real status code, which was 307 here. Point the healthcheck at a route that returns 200 directly, like a dedicated /healthz. If you saw this same failure on a deploy before Aug 28 afternoon UTC, tell us, that would be something different.


Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...