Edge in us-east4 resets every WebSocket to my service in synchronized waves (code 1006), upstream healthy
maurier
PROOP

17 days ago

Project: RDomino

Domain: play.rdomino.com -> wotrrh1n.up.railway.app (69.46.46.89). Edge region seen in HTTP logs: us-east4-eqdc4a. Response header: server: railway-hikari.

App: Node/NestJS + Socket.IO (WebSocket, a few hundred concurrent long-lived connections, real-time multiplayer game).

WHAT HAPPENS

Several times a day, every WebSocket connection to the service is cut at the same moment. Each "episode" is 8 waves: two waves about 60 s apart, repeated every ~5 minutes, 4 times, then it stops. In each wave the whole connected population (same client IPs in both waves of a pair) is closed and all of them reconnect within 1-2 s.

Episodes so far (UTC):

  • 2026-09-17 13:43-13:59
  • 2026-09-18 13:09-13:25, 15:16-15:32, 16:19-16:35, 17:04-17:20, 17:49-17:5x

Exact wave starts on the current deployment (client reconnects seen by my server): 17:04:46, 17:05:44, 17:10:18, 17:11:16, 17:15:14, 17:16:18, 17:19:25, 17:20:31, 17:49:10, 17:50:21.

EVIDENCE THAT IT IS NOT MY SERVICE

  • The container never restarts and the process never stalls: one boot per deployment, CPU 0.1-0.5 vCPU, memory 1.1-1.4 GB, and the stdout log stream has no gap longer than 6 s right before any wave. The edge HTTP logs show no 5xx from my service in these windows.

  • Ordering: my server logs the NEW connections first and only ~12 s later sees the old sockets close with Socket.IO reason "transport close" (never "ping timeout", never a client-initiated close). Your HTTP logs show the same: the WebSocket rows ("connection upgraded to WebSocket", status 0) for the old connections are stamped ~12 s after the clients had already reconnected, with lifetimes of up to 25 minutes. Example rows from the 17:04:46 wave, all ended 17:05:01-17:05:02 UTC:

    ZUI_WeWtTISXsxoyBT7zVQ (lifetime 1523 s)

    E78uxLRFTB-JhBRzozsQ6Q (1038 s)

    NxwjsapwQO-IjFTQBhdwDg (808 s)

    9RshzGgvR56OiizpwoOzXw (1523 s)

    So the client side of each connection dies first, the client reconnects, and the edge closes the upstream side later.

  • Independent probes: I kept three test clients open from an office network no player uses (two plain socket.io-client processes, one with transport=websocket only, one default; plus a headless Chromium). At 17:49:30.15 UTC all three were closed in the same second with WebSocket close code 1006 (abnormal closure, no close frame received), 15.2 s after the last server ping (so one ping was already missing). They reconnected in ~1.5 s and were cut again at 17:50:21 and 17:50:37 (1006 again), matching the second wave the players got. A plain client with no application code being cut at the same instant as the players rules out my client and the players' carriers.

  • It affects clients on every carrier and arrives through all of your proxy replicas at once, so it is not one proxy replica.

HOW TO REPRODUCE

Keep any WebSocket open to wss://play.rdomino.com/socket.io/?EIO=4&transport=websocket (or npx wscat -c on that URL) and wait for the next episode; it will close with 1006 while the upstream never closes it. With Node:

const { io } = require('socket.io-client');

const s = io('https://play.rdomino.com', { transports: ['websocket'], auth: { guestId: 'probe-' + Date.now(), name: 'probe' } });

s.on('connect', () => console.log(new Date().toISOString(), 'connect', s.id));

s.on('disconnect', (r, d) => console.log(new Date().toISOString(), 'disconnect', r, d && d.description, d && d.context && d.context.code));

QUESTIONS

  1. Is the edge/proxy fleet serving us-east4-eqdc4a restarting, draining or re-balancing at those timestamps? The 2-waves-per-minute, every-5-minutes, 4-times shape looks like a rolling operation.
  2. Can you see, on your side, what closed the downstream TCP connections at 17:49:30 UTC (and 17:05:00, 17:10:30, 17:15:14) for the request IDs above?
  3. Is there any WebSocket idle/lifetime/connection-count policy on the edge that could do this? My connections carry a Socket.IO ping every 10 s.
  4. If this is a known edge issue, can the service be pinned to a different edge region until it is fixed?

IMPACT

Each wave freezes every live game for up to 20 s (the dead socket is only detected by the heartbeat) and then forces a full resync for every connected player; players are reporting halts several times per hour.

$20 Bounty

4 Replies

Railway
BOT

17 days ago

We've looked into this from our side and haven't found anything on the Railway platform that explains what you're seeing, so working it out means digging into your specific setup.

That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway • 17 days ago


Railway

We've looked into this from our side and haven't found anything on the Railway platform that explains what you're seeing, so working it out means digging into your specific setup. That's exactly what the Railway community is good at, so we'd like to open your thread as a community [bounty](https://docs.railway.com/community/bounties). Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly. **Opening it makes this entire thread public**, including everything already posted. Nothing becomes public until you decide. Use the buttons below. - **Open to the community** - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away. - **Keep it private and close the thread** - Nothing becomes public and the thread closes.

maurier
PROOP

17 days ago

Before opening this up: "nothing on the platform" doesn't match what independent clients see. Three probes on a network no player uses were closed with WebSocket 1006 (no close frame) in the same second as ~300 players, at 17:49:30 UTC, and again at 17:50:21 and 17:50:37, while my container never closed them (its logs show the closes arriving ~12 s later as "transport close") and never restarted. Could someone look at what terminated the downstream TCP side of request ZUI_WeWtTISXsxoyBT7zVQ (ended 17:05:01 UTC) and the connections ending 17:49:30 UTC on edge us-east4-eqdc4a? That is the one thing only your side can see.


Status changed to Awaiting Railway Response Railway • 17 days ago


Railway
BOT

17 days ago

The community is still the best next step here. The buttons below are still live: open the thread up as a public bounty after editing out anything sensitive, or keep it private and close it.


Status changed to Awaiting User Response Railway • 17 days ago


Railway
BOT

17 days ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway • 17 days ago


maurier
PROOP

17 days ago

This is not isolated to my service. https://station.railway.com/questions/websocket-connections-closing-with-code-2c5dd265 (Sept 14, AMS region, unrelated project) describes the identical signature: code 1006 on every connected client simultaneously, every few minutes, Node process never restarting, reconnects working immediately. The community suggestions there (Cloudflare double-proxy, healthcheck timeouts, multiple replicas, missing heartbeat) don't apply to my service: direct CNAME to Railway with no Cloudflare, no healthcheck configured, a single replica, and a 10 s Socket.IO ping that the probes show alive until the cut. Two regions, two customers, same week, same pattern points at the edge.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...