Healthcheck never reaches a confirmed-healthy container: app responds 200 internally, deploy still marked unhealthy
louisdurden
PROOP

14 days ago

Service: forma (project Forma). Multiple deploys today failed their healthcheck (path /api/health, 180s timeout) and were rolled back, even though I confirmed via railway ssh that the app was fully healthy and serving correctly the whole time.

Evidence: Deployment 3bf9fe7a-d4d7-4d2c-8db8-f8a3a5cb8269 (commit 7db9d23) and several others today: build succeeds, container starts, Starting Healthcheck begins, all 8 retry attempts over 3m0s report service unavailable, deploy marked FAILED and rolled back to the previous instance.

Right after a fresh deploy entered the DEPLOYING state (before Railway killed it), I ran railway ssh into the live instance and inspected it directly. next-server (PID 276) was running, state S (sleeping), not crash-looping. It WAS listening: /proc/net/tcp6 showed the wildcard address on port 1F90 hex (8080 decimal) in state 0A (LISTEN), a dual-stack IPv6 socket. It was NOT visible in /proc/net/tcp (the IPv4 table) at all.

I then ran a small Node script inside the container hitting the health endpoint on both loopback addresses, port 8080, path /api/health. IPv4 loopback responded 200 with body status ok, db ok, latencyMs 14. IPv6 loopback responded 200 with body status ok, db ok, latencyMs 16. So the app is unambiguously healthy and responding correctly on loopback from inside the container, at the exact same time Railway own healthcheck prober reports every attempt as service unavailable.

This happened across several different configurations of the service today (Dockerfile builder and Railpack builder, npx versus direct node_modules bin binary invocation, with and without extra env vars), ruling out application code as the cause. The one constant is that it only ever happens on a real orchestrated Railway deploy. The exact same command, image and env, run via docker run locally or via railway ssh and railway run, never reproduces any hang.

There is a closely related thread from 6 months ago on this forum, titled Deployment healthcheck fails even though candidate instance returns 200 in-container, with the same exact symptom and the same SSH-based verification method. A reply there suggested this is likely the healthcheck prober failing to reach the container over its actual private-network IP as opposed to loopback, possibly tied to legacy IPv6-only private networking environments not yet on the IPv4 rollout described in the Support IPv4 Private Networks thread. That matches exactly what I found: loopback works fine from inside the container, but the app was only ever visible listening in the IPv6 socket table, never in the IPv4 one.

Could someone confirm whether this project environment (Forma, service forma) is still on legacy IPv6-only private networking, and if so migrate it, or point me at how to do so myself? Happy to provide more deployment IDs, exact timestamps, or run further diagnostics. I have SSH access to a live instance and can capture whatever is needed on request.

Awaiting Conductor Response$20 Bounty

3 Replies

Railway
BOT

14 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 14 days ago


admilucky1-cyber
FREE

14 days ago

The behavior you are encountering is a classic issue in dual-stack and IPv6-enabled container networking environments.

The Root Cause

Your troubleshooting via /proc/net/ revealed the key detail:

  • When Next.js or Node.js binds to port 8080 without an explicit IPv4 bind address, it creates an AF_INET6 socket listening on [::]:8080.
  • On dual-stack Linux systems with IPV6_V6ONLY=0 (the kernel default), an AF_INET6 wildcard listener accepts connections from both IPv6 and IPv4 (via IPv4-mapped IPv6 addresses ::ffff:x.x.x.x).
  • This is why testing locally on 127.0.0.1 and ::1 inside the container via railway ssh returns 200 OK—the Linux loopback interface seamlessly maps IPv4 loopback requests into the IPv6 dual-stack socket.
  • However, Railway’s internal healthcheck prober routes traffic to the container’s assigned Private IPv4 Address or Private IPv6 Hostname/Address over the internal mesh overlay network.

In legacy IPv6 or dual-stack environments, if the orchestrator probe attempts an IPv4 connection over the container's eth0 network interface (e.g., hitting 10.x.x.x:8080), Linux kernel netfilter/routing rules or network namespace boundaries on dual-stack sockets often drop or reject incoming external IPv4 packets targeting an AF_INET6 socket if the application didn't explicitly bind an AF_INET (0.0.0.0) listener.

Solutions

While you wait for Railway support to check and migrate your project’s network configuration, you can resolve this immediately at the application or framework level.

Option 1: Explicitly Bind Next.js / Node to 0.0.0.0

Force your application process to explicitly open an AF_INET socket on 0.0.0.0 rather than relying on default dual-stack socket binding.

  • If launching via Next.js CLI:

    Update your startup script in package.json or Procfile/Dockerfile CMD:

    next start -H 0.0.0.0 -p 8080

  • If using Environment Variables:

    Set HOSTNAME=0.0.0.0 and PORT=8080 in your Railway Environment Variables so the Next.js standalone server explicitly binds to IPv4 0.0.0.0.

Option 2: Bind Dual Sockets explicitly (if IPv6 private networking is strictly required)

If your Node process uses a custom server or HTTP listener (e.g., Express/Fastify), ensure it binds explicitly to 0.0.0.0 or creates dual listeners for 0.0.0.0 and ::.

// Force explicit IPv4 listen

server.listen(8080, '0.0.0.0', () => {

console.log('Listening on 0.0.0.0:8080');

});

Verification Step

Once you apply HOSTNAME=0.0.0.0 or next start -H 0.0.0.0:

  • Trigger a deploy and railway ssh into the candidate instance.
  • Check /proc/net/tcp (IPv4 table). You should now see entry 00000000:1F90 in state 0A (LISTEN).
  • The Railway healthcheck prober will immediately succeed on its next attempt over the private IPv4 interface.

bl4ckph4ntom1
PROTop 10% Contributor

13 days ago

I don't think I'd switch this to 0.0.0.0 yet.

The [::]:8080 listener by itself isn't evidence that IPv4 is broken. On Linux, if net.ipv6.bindv6only=0, that one IPv6 socket can also accept IPv4-mapped connections, which seems to line up with your 127.0.0.1:8080 test already returning 200.

You can confirm that quickly with:

cat /proc/sys/net/ipv6/bindv6only

If it returns 0, then the missing entry in /proc/net/tcp isn't really the smoking gun here.

I'd check two things before changing the bind address.

First, reproduce the request with the exact Host header Railway uses:

curl -i --max-redirs 0 \
  -H "Host: healthcheck.railway.app" \
  http://127.0.0.1:$PORT/api/health

Railway's healthcheck uses healthcheck.railway.app, and "service unavailable" isn't always specific enough to tell whether this was a connection failure vs an HTTP response the prober didn't accept.

Second, check the container's actual non-loopback addresses:

ip -br addr

and try /api/health against the container's own private IPv4/IPv6 address, not just loopback.

If the Host-header test still gives a direct 200, and the endpoint also works on the non-loopback private address, then I think you've ruled out most of the application/binding side pretty convincingly.

Also, if this really is a legacy environment, Railway's docs say those environments resolve private networking over IPv6 only. That actually makes [::] the safer bind here, not 0.0.0.0.

So if all of the above passes but Railway's probe still never reaches the app, I'd lean toward the candidate deployment not being attached/routable to the healthcheck path correctly rather than a Next.js listener problem.

At that point the deployment ID + exact healthcheck timestamps are probably what Railway needs to inspect on their side.


louisdurden
PROOP

13 days ago

Ran both checks you suggested, directly on the live container via railway ssh:

cat /proc/sys/net/ipv6/bindv6only -> 0. You're right that this alone doesn't prove IPv4 is broken.

But: ls /sys/class/net shows lo, railnet0, tunl0 - a real non-loopback network interface. hostname -I returns 10.135.245.133 (real private IPv4) plus an fd12:... IPv6 address. So this environment is genuinely dual-stack, not legacy IPv6-only.

I then hit the health endpoint on that real private IPv4 address, not loopback (curl wasn't in the slim image, used python3 urllib instead): http://10.135.245.133:8080/api/health returned 200 {"status":"ok","db":"ok"}. That's the exact test you asked for, and it passes.

Applied next start -H 0.0.0.0 anyway and deployed. Post-deploy, /proc/net/tcp on the same container now shows 00000000:1F90 in state 0A (LISTEN) and /proc/net/tcp6 is empty - so before the fix it really was IPv6-only despite the dual-stack env, and Railway's healthcheck prober (which was failing every attempt) now passes cleanly. Deploy is green and has stayed healthy since.

So both things were true at once: real dual-stack networking (your point), and a genuinely IPv6-only listener that the healthcheck prober couldn't reach over the private IPv4 path (the original fix). Didn't test the Host-header variant since the direct private-IP test already settled it. Thanks for pushing on this - the extra verification was worth doing before shipping it.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...