a month ago
Subject: Persistent request timeouts / dial failures on production service, unrelated to app code — 7+ hours downtime
Project: bglam-stock-transferProject ID: f93ce332-fd77-4fb6-b7cd-d4494c7186f5Service: bglam-stock-transferService ID: ae1442af-f7f7-48ff-9c0c-0b2c5a3332c2Environment: production (726ee7d9-53d4-4720-914c-cdd8589b0379)Region: sfoPublic domain: bglam-stock-transfer-production.up.railway.app
Summary
Since approximately 2026-08-28 05:21 UTC, our production service has been intermittently or fully unable to respond to inbound HTTP requests (both browser page loads and inbound Shopify webhooks) within a reasonable window. Requests hang for 5-40+ seconds and are then terminated as HTTP 499 (client gave up) or 502, with upstreamErrors in the HTTP logs reporting "connection dial timeout" and "client has closed the request before the server could send a response".
This is still ongoing as of this writing (~12:30 UTC same day — 7+ hours). It is blocking our warehouse staff from using the application at all.
We have spent several hours ruling out application-level causes and would like Railway's help determining whether this is a platform-side networking/proxy issue affecting this service.
What we've ruled out
- Not a code/version issue. We have deployed and tested versions 103.181.0 through 103.187.0 today — including versions that were running cleanly before 05:21 UTC — and every single one exhibits the identical failure signature once live. Rolling back further has not helped.
- Not (solely) an application-side concurrency issue. We identified that our app fires several background sync jobs at boot through a shared concurrency limiter, and disabled them via env vars to rule this out. This measurably improved
/api/health(now returns 200, though still slow at 5-11s), but general page loads (GET /) and most inbound webhooks still fail the same way. So while our own job scheduling was making things worse, it is not the root cause. - Railway status page (status.railway.com) shows no declared incident at time of writing, but we understand isolated/regional issues may not appear there.
Concrete evidence (deployment instance IDs + timestamps)
Deployment 111e7d8a-8302-4b2d-932f-a47a28c02808 (instance 5f9392fe-b47b-4554-b944-eadd7b433e13), live 2026-08-28T00:19:41Z through ~06:46Z:
- Ran cleanly (HTTP 200) from ~01:48 UTC through ~05:00 UTC.
- From ~05:20 UTC onward, degraded sharply: of 501 sampled HTTP log entries, 389 returned 499 and 15 returned 502.
upstreamErrorsrepeatedly showed entries like:[{"deploymentInstanceID":"5f9392fe-b47b-4554-b944-eadd7b433e13","error":"connection dial timeout","duration":5000},{"deploymentInstanceID":"5f9392fe-b47b-4554-b944-eadd7b433e13","error":"client has closed the request before the server could send a response","duration":939}]
Deployment 4f5283bf-8772-4602-9238-dbde79369aaf (instance 64fe7d87-29a0-441f-9207-b38061dff780), booted 2026-08-28T12:14:45Z:
- 100% of sampled requests — both
POST /webhook/shopify/*and plainGET /from a real browser — returned 499 from the very first request onward, continuously, for at least 8 minutes straight (12:14:45 - 12:22:35 UTC observed). - Example:
GET /at 12:20:39Z took 10,472ms before the client gave up; another at 12:21:19Z took 39,690ms.
Deployment 545ca020-e4ee-4c0f-9be5-7a2653818fc3 (instance d2f5dfe9-bcf1-486f-b2d2-5ce22c67fc92), booted 2026-08-28T12:29:19Z, after we disabled the 5 background sync jobs mentioned above:
/api/healthnow returns 200, but slowly (5-11 seconds per request, tested repeatedly).GET /still times out entirely — two consecutive test requests at 12:31:10Z and 12:31:12Z took 19,677ms and 14,908ms respectively before failing with the same"client has closed the request before the server could send a response"error.- Most
POST /webhook/shopify/*requests still return 499, though a small number succeed quickly (e.g. one in 30ms).
What we'd like help with
- Please check whether there is platform-side degradation (proxy/dial layer, networking, or compute-node resource contention) affecting this specific service or the
sforegion starting around 2026-08-28 05:21 UTC. - We don't have access to the
upstreamRqDurationfield orX-Railway-Request-Idresponse headers through our current tooling to distinguish network-path latency from application response time ourselves — if you're able to pull that server-side for the deployment instance IDs above, that would help confirm where the delay is actually occurring. - Given this has now persisted across 6+ different deployments/redeploys and both freshly-built and reused cached images, we don't believe further app-side changes on our end will resolve it, but we're glad to provide anything else useful.
This has caused 7+ hours of downtime for a production system our warehouse staff depend on to work. We'd appreciate urgent attention.
Thank you.
2 Replies
a month ago
Having looked into this, the issue appears to be in your application code or configuration rather than the Railway platform itself, which puts it outside what Railway support can resolve directly.
This is exactly the kind of problem the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.
Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.
- Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
- Keep it private and close the thread - Nothing becomes public and the thread closes.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • about 1 month ago
a month ago
Thanks for looking into this. Respectfully, I don't think that conclusion holds up against what we've now observed — sharing the fuller evidence here now that the thread's open.
The core problem with "it's your code": we reproduced the identical failure on byte-for-byte unchanged code that had already been running stable in production for hours.
Timeline from yesterday (2026-08-28) through today:
- A new code change was deployed and failed Railway's healthcheck. Redeploying the identical build failed identically.
- We manually reverted the code to the exact previous form — the same code that had already run stable in production for 5+ hours before this incident — and redeployed. It failed the same way, five more times in a row.
- On one of those reverted-code failures, Railway's own railway-agent diagnostic tool confirmed our container had already finished booting and was actively accepting connections on its port for 3+ minutes, while the healthcheck prober still reported "service unavailable" with zero HTTP requests ever reaching the container.
- Production came back on its own roughly 80 minutes later, on an older build we deployed from a local zip — again, no application code explains why that specific redeploy succeeded when nothing else had.
- Later the same day, we repeated this entire experiment independently: a different code change (a UI text/nav rename, nothing touching startup, networking, or health-check logic) was deployed and failed. We reverted to the exact prior stable code and redeployed — it failed identically, again. Total across both incidents: 11+ deploy attempts, at least four distinct code states (two genuinely different app versions, plus the same "known-good" code redeployed multiple times), every failure with the same signature, and the outcome never once correlated with which code was actually running.
- Today (08-29), we attempted one more deploy of our fully-tested, previously-certified stable build (103.187.0). This time the failure mode was different again: the deployment built successfully (image pushed, confirmed in build logs) but then sat in "Building" status for 20+ minutes with zero deploy-log output and zero progress toward the health-check phase at all — not a healthcheck timeout this time, just no forward progress. No code was involved in that stall; the exact same source had already built and deployed cleanly hours earlier.
If this were genuinely an application code or configuration issue, we would expect the reverted, previously-stable build to succeed when redeployed. It didn't — repeatedly, across two separate incidents on two separate days, with the failure taking multiple different shapes (healthcheck timeout with a confirmed-listening container, and now a build that completes but the deployment never advances). The one constant across all of this is the platform-side deploy/proxy layer, not our code.
We're not asking you to take our word for it — the deployment IDs and timestamps for all of this are in our original message and above. We'd genuinely appreciate someone taking a second look with this fuller picture, rather than closing this out as application-side. Happy to provide anything else that would help — logs, deployment IDs, timing windows, whatever's useful.
For now we're not going to keep hammering redeploys against what looks like a platform issue — we'll hold at our last known-good build and wait to hear back.
Thank you.