a month ago
Hi,
I'm troubleshooting a Railway staging deployment where the application builds and starts successfully, but the deployment consistently fails during the Railway healthcheck.
Environment:
- Next.js 16.3.0
- Node.js 24.19.0 on Railway
- Railpack
- PostgreSQL Railway service
- Staging environment
- Healthcheck path: /api/health
- Healthcheck timeout: 300 seconds
- Application uses Railway's injected PORT
Deployment:
- Initialization: PASS
- Build: PASS
- Pre-deploy: PASS
npx prisma migrate deploy: PASS- 14 migrations found, no pending migrations
npm start/next start: PASS- Next reports Ready in ~200 ms
- Local address: localhost:8080
- Network address: private IPv4:8080
- Network > Healthcheck: FAIL after approximately 4m30
The healthcheck endpoint is an unauthenticated App Router GET route at:
/api/health
It returns HTTP 200 when the same production build and staging configuration are reproduced outside the Railway deployment healthcheck.
The endpoint does not need PostgreSQL, R2, Stripe, Resend, or any other external network call to return its healthy response with the current staging flags.
We also verified:
- no middleware/proxy intercepts /api/health
- no redirect or authentication guard applies to it
- Next.js uses Railway's PORT
- the process does not show an explicit crash after "Ready"
- increasing healthcheckTimeout from 30 to 300 was applied correctly, but only extended the failure duration
- notification, SMS and payment transports are explicitly disabled
- the notification configuration parser was independently reproduced with the same non-secret staging values and passes
- /api/health returns HTTP 200 in a real local
next startproduction-build test - Railway Network Logs contain no useful healthcheck request information
Railway's automatic Diagnose suggested that notification environment variables might cause a silent crash. We audited and reproduced that configuration independently and confirmed that this is not the case: the parser passes and cannot terminate the Next.js process.
One complication: the Railway service Console appears to connect to the previous ACTIVE deployment rather than the failed candidate deployment, so probing localhost from that Console does not test the candidate container.
Could you help determine:
- Whether Railway's deployment healthcheck is actually reaching the candidate container on the injected PORT.
- Whether there is a way to inspect/probe the candidate container before healthcheck promotion.
- Whether there are known issues with Railway healthchecks and Next.js 16 / Railpack in this situation.
- Whether the healthcheck request Host header, candidate networking, or deployment routing could explain this behavior.
I would prefer to identify the actual cause before disabling the healthcheck or changing application code.
Thanks.
5 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 2 months ago
a month ago
Try removing the healthcheck path, and check if the app is accessible from its public URL. If it's not, that means you either have a misconfigured hostname or port. Since you are listening on the PORT env variable, i'd rule that one out. So, I recommend you set HOSTNAME=0.0.0.0 in your service variables, just in case next.js is binding to localhost.
darseen
Try removing the healthcheck path, and check if the app is accessible from its public URL. If it's not, that means you either have a misconfigured hostname or port. Since you are listening on the PORT env variable, i'd rule that one out. So, I recommend you set `HOSTNAME=0.0.0.0` in your service variables, just in case next.js is binding to localhost.
a month ago
Thanks. I checked the binding side before posting.
The candidate deployment logs show:
- Next.js 16.3.0
- PORT=8080 provided by Railway
- Local: http://localhost:8080
- Network: http://:8080
- Ready in ~200 ms
Next.js 16 documents 0.0.0.0 as the default hostname for next start, so I haven't added HOSTNAME yet because I wanted to avoid changing variables without evidence.
I also reproduced the exact production build locally with the same non-secret staging flags and /api/health returns HTTP 200.
The difficult part is that Railway's Console appears to attach to the previous ACTIVE deployment, not the failed candidate deployment, so I cannot probe localhost/private IPv4 inside the candidate while it is waiting for healthcheck promotion.
Would you still recommend setting HOSTNAME=0.0.0.0 first, or temporarily removing only the healthcheck path to let the candidate become ACTIVE and then probe /api/health directly from that exact deployment?
I'm trying to use the smallest diagnostic change possible.
a month ago
I'd go with removing the healthcheck temporarily, just to see if the app is working and binding correctly. If you can share the URL, it would be helpful as well.
a month ago
Hi lnxbeats — building on darseen's suggestion, which I think is very good and points in the right direction. I'd just like to add some evidence I found on why it matters and one small adjustment on how to apply it, in case it helps you make the smallest possible change.
-
Railway's healthcheck probe is IPv4. Railway staff have stated this in another Central Station thread ("health checks use ipv4"): https://station.railway.com/questions/fast-api-service-health-check-fails-in-ip-a0add1f5 — and in that thread, as well as in https://station.railway.com/questions/uvicorn-health-check-fails-with-i-pv6-bin-6e9f929e, services that were listening on :: only had exactly your symptom: process reports ready, no crash, no request ever logged, healthcheck times out.
-
What next start binds to when no -H is given. Looking at the Next.js source (packages/next/src/server/lib/start-server.ts), when no hostname is passed the server does server.listen(port, undefined), which lets Node pick the address — and Node's default in that case is :: (IPv6 any), not 0.0.0.0. The --help text says "default: 0.0.0.0", but there is an open/closed discrepancy issue about that: https://github.com/vercel/next.js/issues/68836. Your log line Network: http://:8080 with an empty host is consistent with this: Next only fills that host when it finds a global IPv6 interface (lib/get-network-host.ts), and inside the Railway container it didn't find one.
-
The small adjustment: use the -H flag, not the HOSTNAME env var. As far as I can tell from the CLI definition (packages/next/src/bin/next.ts), -p/--port is declared with .env('PORT'), which is why PORT from Railway works — but -H/--hostname has no .env(...) binding, so next start does not read a HOSTNAME variable. (HOSTNAME is only read by the server.js generated in output: 'standalone' mode, which isn't your setup with Railpack.) So setting HOSTNAME=0.0.0.0 as a service variable may not change anything here, while this should:
next start -H 0.0.0.0 -p $PORT
as your start command (or "start": "next start -H 0.0.0.0 -p $PORT" in package.json).
- How to verify without guessing. After the change, the first "Ready" log should show Network: http://0.0.0.0:8080 (or the container's IPv4) instead of an empty host. If the healthcheck then passes, that confirms the probe simply wasn't able to reach an IPv6-only listener.
Also, unrelated to the bind but worth keeping in mind: Railway sends the healthcheck with the hostname healthcheck.railway.app (https://docs.railway.com/guides/healthchecks), so if you ever add host-based restrictions (allowedHosts, a middleware checking Host), that name needs to be allowed.
I may be missing something about your exact setup, so please treat this as a hypothesis with supporting evidence rather than a certainty — but it lines up well with every detail you described (ready in 200 ms, no crash, local build passing, nothing in Network Logs). Hope it helps.
rluf
Hi lnxbeats — building on darseen's suggestion, which I think is very good and points in the right direction. I'd just like to add some evidence I found on why it matters and one small adjustment on how to apply it, in case it helps you make the smallest possible change. 1. Railway's healthcheck probe is IPv4. Railway staff have stated this in another Central Station thread ("health checks use ipv4"): https://station.railway.com/questions/fast-api-service-health-check-fails-in-ip-a0add1f5 — and in that thread, as well as in https://station.railway.com/questions/uvicorn-health-check-fails-with-i-pv6-bin-6e9f929e, services that were listening on :: only had exactly your symptom: process reports ready, no crash, no request ever logged, healthcheck times out. 2. What next start binds to when no -H is given. Looking at the Next.js source (packages/next/src/server/lib/start-server.ts), when no hostname is passed the server does server.listen(port, undefined), which lets Node pick the address — and Node's default in that case is :: (IPv6 any), not 0.0.0.0. The --help text says "default: 0.0.0.0", but there is an open/closed discrepancy issue about that: https://github.com/vercel/next.js/issues/68836. Your log line Network: http://:8080 with an empty host is consistent with this: Next only fills that host when it finds a global IPv6 interface (lib/get-network-host.ts), and inside the Railway container it didn't find one. 3. The small adjustment: use the -H flag, not the HOSTNAME env var. As far as I can tell from the CLI definition (packages/next/src/bin/next.ts), -p/--port is declared with .env('PORT'), which is why PORT from Railway works — but -H/--hostname has no .env(...) binding, so next start does not read a HOSTNAME variable. (HOSTNAME is only read by the server.js generated in output: 'standalone' mode, which isn't your setup with Railpack.) So setting HOSTNAME=0.0.0.0 as a service variable may not change anything here, while this should: next start -H 0.0.0.0 -p $PORT as your start command (or "start": "next start -H 0.0.0.0 -p $PORT" in package.json). 4. How to verify without guessing. After the change, the first "Ready" log should show Network: http://0.0.0.0:8080 (or the container's IPv4) instead of an empty host. If the healthcheck then passes, that confirms the probe simply wasn't able to reach an IPv6-only listener. Also, unrelated to the bind but worth keeping in mind: Railway sends the healthcheck with the hostname healthcheck.railway.app (https://docs.railway.com/guides/healthchecks), so if you ever add host-based restrictions (allowedHosts, a middleware checking Host), that name needs to be allowed. I may be missing something about your exact setup, so please treat this as a hypothesis with supporting evidence rather than a certainty — but it lines up well with every detail you described (ready in 200 ms, no crash, local build passing, nothing in Network Logs). Hope it helps.
a month ago
Thanks for the detailed analysis. We managed to identify the root cause and confirm it with direct probes from the actual candidate container.
The binding was not the issue.
Once we temporarily removed the Railway healthcheck and allowed the candidate deployment to become ACTIVE, we probed /api/health directly using both:
- 127.0.0.1:$PORT
- the container private IPv4:$PORT
Both interfaces were reachable, but both returned:
HTTP 503
check=media-storage
We then diagnosed the media configuration and found that three R2 values were not actually present in the runtime container:
- MEDIA_S3_ENDPOINT
- MEDIA_S3_ACCESS_KEY_ID
- MEDIA_S3_SECRET_ACCESS_KEY
After correctly injecting those values, the same probes returned:
loopback HTTP 200 ok=true
private HTTP 200 ok=true
We then restored the Railway healthcheck at /api/health, and the deployment passed successfully and became ACTIVE with the existing next start command.
So in this case the healthcheck/networking behavior was correct: Railway was rejecting the application's legitimate 503 response.
Thanks again — the IPv4/IPv6 analysis was useful for ruling out the binding layer.