18 days ago
Hi Railway,
We're seeing long-lived WebSocket connections to our service get cut in recurring bursts. Many unrelated clients are affected within the same second or so, and we'd like help finding the cause.
Setup
- Service
cbffacbd-7d9d-4abb-adc7-e99f172b59b6, deployment29364eaf-d5fd-4efd-be17-9ff76a89d3d0. - Region europe-west4-drams3a, 1 replica. Custom domain (happy to share privately), with no Cloudflare or other proxy in front.
- Node.js with socket.io 4.x. WebSocket transport only, no long-polling.
- Engine.io defaults: the server pings every 25 s and waits up to 20 s for the pong.
- Clients are Chromium-based browsers on home internet connections in Poland, one connection each, open for hours.
What we observe
A burst lasts about 15-30 minutes. During it, connections from different clients and different ISPs are dropped together:
- The client sees the connection close abruptly (close code 1006, no close frame) and reconnects about 2 s later.
- Our server is not told. The server-side socket stays open until our next ping goes unanswered. Then the connection closes as
transport close, about 15 s after the undelivered ping. Your HTTP logs show the same delay: the entry for the upgraded request ends 15-25 s after the client saw the close. - The new connections are often cut again within about a minute (23 of 36 reconnects on our control connection, below): either 15 s after connecting, exactly when a ping is sent (25 or 50 s), or 15 s after a ping (40 or 65 s). So a connection is cut either at a data exchange or about 15 s after one. This repeats for the length of the burst, then everything is stable for roughly an hour or more.
Our process shows no errors, no event-loop stalls and no restarts during bursts. Pings to connections that are not cut keep flowing normally.
A control connection outside our user base
To rule out our clients' networks, we ran a minimal socket.io client on a separate home connection (different ISP) that only answers pings.
- Cut in the same bursts: it was cut 34 times between 2026-09-16 22:40 and 2026-09-17 05:05 UTC, in four bursts that overlapped the bursts on our real clients.
- Healthy until each cut: pings arrived every 25 s right up to each cut; the longest gap was 25.9 s.
- At low load: some bursts ran at 3-5 AM with only 3-5 connections open, so they don't depend on load.
- Same country: our clients and the control connection are all in Poland. We haven't yet tested from another country, so a problem on the network path from Poland to your edge isn't ruled out.
Burst times (UTC)
| Burst | Clients affected | Connections cut (a client can be cut several times) |
| 2026-09-16 13:25 | 7 | 7 |
| 2026-09-16 20:47-21:14 | 14 | 129 |
| 2026-09-16 22:40-23:05 | 5 | 46 |
| 2026-09-16 23:59-00:23 | 5 | 37 |
| 2026-09-17 02:56-03:21 | 3 | 25 |
| 2026-09-17 04:38-05:04 | 3 | 29 |
Counts come from our server's socket-disconnect log and do not include the control connection below.
Also 2026-09-15 19:45-20:02, before the current deployment: the control connection was cut 8 times.
Examples with request ids
2026-09-16 20:48:28-20:48:31 UTC. 14 connections that had all opened at 20:16:56-57 (clients reconnecting after a deploy) ended within 3 s of each other. Their request-id suffixes differ, so this looks like many edge instances rather than one:
kpMDLp4HShCdZtkx21mRUA QsMdrzICQUu_APFX7fhULg flSAGACKQfyehuuw8Hk9cQ
cbe8JG9WR4m0zvdLw9P4nw wNdkOmjCQuG7v1_L55So4g pfWLPAI9SMWmp7mvnTga3g
UpZLGrr4TRmGh159ALIuDA EbFFefJAQlarnagznTga3g 9gfHesmLR2GmLIhuALIuDA
k0PNd9zEQaiMP_7qss7a6g tDBGLc-MR86SMtffs_GTAg -XMI4FiDT9661SFYWVMv1w
U9Y3nqP6T4C2nXsavEOTxA Dwh65bZ8Rb6PURTvEuroCgThe control connection in the same burst: gSNylOKmRK-1JiYujUJq2g. The client saw the close at 20:47:57.3; your log entry ends at 20:48:12.2.
2026-09-16 21:13:27-21:13:41 UTC. Connections opened at 21:12:22-36 (reconnects from the previous cut) were cut again after about 65 s:
v9J_YPbhQQemim-L55So4g a_5JwatNTy6svFVQnTga3g Mr7ZnOEbTSObAcE25nX1uw
ti0aQzaYSCSDxFl3vEOTxA I5tPIdgCTdCJp73621mRUA 838fi7e5Ru2gHGJmN8N_Fg
1JuMliEpTcObtbsJ8Hk9cQ BQzXLFOyQg2Dt5ZpN8N_FgThe control connection: xZr_FXP8QD2P4CEs21mRUA.
What we're asking
- Were there edge or proxy events (restarts, rebalancing, maintenance, network incidents) in or near europe-west4-drams3a at these times?
- Can you see why these connections were reset on the client side without the close being passed to the upstream?
- Is there anything on our side (config, a different way to expose the service, TCP proxy vs HTTP) you'd recommend for long-lived WebSockets?
We can give more timestamps, request ids or the control client's logs if useful.
Thanks!
2 Replies
18 days ago
We've looked into this from our side and haven't found anything on the Railway platform that explains what you're seeing, so working it out means digging into your specific setup.
That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.
Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.
- Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
- Keep it private and close the thread - Nothing becomes public and the thread closes.
Status changed to Awaiting User Response Railway • 18 days ago
18 days ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • 18 days ago
16 days ago
Two other threads describe the same fault, and one of them observed the same hours we did:
- https://station.railway.com/questions/websocket-connections-closing-with-code-2c5dd265 (AMS, Node, two projects cut together)
- https://station.railway.com/questions/websocket-clients-are-reset-on-a-15-seco-962605a3 (us-west2, ~12 clients over 8 IPs and 7 edge locations)
That makes three tenants in three regions with one signature: the client gets 1006 with no close frame, the origin socket stays open with no FIN and no error, and the origin only learns when its next ping goes unanswered — after which the edge closes the upstream 15-25 s later.
Our worst window yet: 2026-09-18, 13:07-18:17 UTC
762 connections cut with that signature, across 25 different channels. They arrive in tight bursts, and the cadence is regular:
| Burst (UTC) | Connections cut |
|---|---|
| 13:12-13:14 | 41 |
| 13:21-13:23 | 38 |
| 13:30-13:31 | 29 |
| 13:37-13:39 | 38 |
| 15:20-15:21 | 39 |
| 15:29-15:30 | 54 |
| 15:37-15:38 | 36 |
| 15:45-15:46 | 31 |
| 16:22-16:24 | 34 |
| 16:31-16:32 | 44 |
| 16:39-16:40 | 46 |
| 16:46-16:48 | 37 |
| 17:08-17:09 | 37 |
| 17:16-17:17 | 35 |
| 17:24-17:25 | 26 |
| 17:31-17:33 | 28 |
| 17:52-17:53 | 38 |
| 18:00-18:02 | 48 |
| 18:08-18:09 | 36 |
| 18:15-18:17 | 42 |
Two things stand out in that table:
- Bursts come in GROUPS OF FOUR, about 8-9 minutes apart, with a 25-35 minute gap between groups. That looks like something cycling rather than load.
- 17:08-18:17 overlaps the 17:10-17:59 window in the us-west2 thread above, in a different region.
Then it stopped. 09-19 (Saturday) has had one such drop all day, against 762 the day before. Both other threads report weekday-only windows too.
One correction to the 15-second grid
That thread finds resets landing on 15-second multiples of connection age. Ours do not, and the difference points at the mechanism. Our server pings every 10 s (engine.io pingInterval), and of 765 resets:
- 58% land within 0.5 s of a 10-second multiple (chance would be 10%)
- 22% land near a 15-second multiple (chance 7%)
The short-lived ones cluster at 10, 20, 40, 50 and 60 s of age. So the connection is not cut by a timer of its own: it is cut when our own ping goes out, whatever interval we use. pingUnansweredMs is 13.0-16.6 s (median 15.0) on every one of the 762.
That matters for the workaround advice: a shorter heartbeat cannot avoid this, and may make it more frequent, because the write is what triggers it.
What we already run, so it can be ruled out
One replica, one region, no Cloudflare proxying in front of the custom domain, socket.io heartbeats, healthcheck 60 s, no restarts or errors in the process during bursts, and ordinary HTTP requests succeeding throughout.
The ask for Railway
Please correlate 2026-09-18 13:00-18:30 UTC for europe-west4-drams3a against edge proxy rolling updates or draining, and say whether the 8-9 minute cadence in the table matches a pod replacement cycle. Two other tenants are reporting the same thing in the same hours, so this is a platform question rather than three application bugs.