Established WebSocket connections destroyed without a close frame
aehoe
PROOP

19 days ago

Established WebSocket connections destroyed without a close frame - two services, origin healthy

ENVIRONMENT

Production: femuridle.com

Test: femur-idle-server-test.up.railway.app

(Project, service and deployment IDs were shared privately with Railway support in this thread; happy to share again with anyone who needs them.)

Both in us-east4-eqdc4a. We reach both via edge gru1 (x-hikari-trace: gru1.jrt2 / gru1.469d). Node.js + "ws".

SYMPTOM

In recurring bursts of 10-20 min, several times a day, established WebSocket connections are destroyed at the transport level with no close frame (clients see 1006). Both services are hit together.

The key asymmetry: the CLIENT sees the connection destroyed while the ORIGIN still holds the socket handle open - it logs no close and only closes later, on its own 120 s application timeout. Our code cannot break the TCP path to a client while keeping its own handle alive.

REFERENCE INCIDENT - 2026-09-16, 20:49:51 to 21:00:59 UTC (production)

20:52:57 - 243 sockets closed in 144 s (gru1: 194)

20:57:47 - 289 sockets closed in 153 s (gru1: 224)

21:01:32 - 256 sockets closed in 121 s (gru1: 186)

Process: single boot, uptime 262,660 s continuous, memory stable, concurrency stable (clients reconnect).

Throughout the window HTTP on the same hostname was unaffected: polling GET /api/online every 5 s gave 138 samples per service, ZERO failures, 134-412 ms. A brand-new WebSocket opened in 304-571 ms in the same millisecond an existing one was destroyed.

ORIGIN WAS NOT DEGRADED WHEN IT STARTED (30 s health samples; p99 is the worst of each period)

20:48 52.1 ms | 20:49 42.3 ms <- burst begins | 20:52 86.0 | 20:56 203.7 | 20:57 476.3 | 21:00 140.4 | 21:01 42.0 <- back to baseline

Latency is at the window MINIMUM when the burst starts and only rises 7 minutes in, after thousands of reconnections are already in flight. It is a consequence, not a cause.

CONTROL EXPERIMENT (our machine, Brazil, via gru1)

4 idle WebSockets (2 per service, with and without permessage-deflate) plus 2 WebSockets to a NON-RAILWAY host, same machine, same link.

During burst (20:54-21:01): Railway connections destroyed 21, average lifetime 45 s. Non-Railway control: 0.

After burst (21:01-21:14): Railway 0. Probe behaviour identical in both phases.

The control stayed up across all 21 destructions.

RULED OUT, BY MEASUREMENT

Origin crash/restart (3+ days uptime); one bad container (two separate services hit together); our code performance (latency at minimum when it starts); permessage-deflate (with and without die alike); our network link (control survives); our probe causing it (zero losses outside bursts); DNS/routing (HTTP 200 and new handshakes succeed at that instant); application-level close (origin logs none).

We also audited our own send path: in a load test with 110 simulated clients in one city, the largest outbound frame is 1,364 B (largest 100 ms burst per client 3,700 B). Session boot is 275 KB and was benchmarked at 110 simultaneous boots - slowest client fully served in 234 ms. Node's requestTimeout/headersTimeout do not reach upgraded sockets. We found no mechanism on our side that produces this signature.

WHAT WE ARE ASKING

For the windows below, on edge gru1 (gru1.jrt2 / gru1.469d) toward us-east4-eqdc4a:

  1. Any event that terminates ESTABLISHED WebSocket/upgraded connections while still accepting new ones and serving HTTP on the same hostname - edge node drain/rotation/redeploy, connection-table limits, upstream connection recycling, or idle/lifetime caps on upgraded connections.
  2. Whether such events are expected to propagate WITHOUT a close frame and to leave the ORIGIN side half-open. The origin is never informed, which is what turns a transient event into a user-visible failure for us.
  3. Whether two distinct services in the same region are expected to be hit by the same edge-side event.

If there is configuration on our side that would make this graceful (a close frame propagated to both ends), we would like to know.

TIME WINDOWS (UTC)

09-14: 11:36-11:50, 12:46-13:00, 14:48-15:02, 16:03-17:03, 18:50-19:05

09-15: 09:44-09:57, 11:01-11:16, 12:29-12:43, 13:50-14:04, 15:01-15:15, 19:44-19:58, 22:29-22:44

09-16: 08:55-09:09, 09:44-09:58, 20:49-21:00

Bursts cluster between 08:00 and 13:00 UTC (46% of all lost connections), which is NOT our peak - concurrency is at its lowest then.

SAMPLE x-railway-request-id (HTTP requests that SUCCEEDED in the same millisecond a WebSocket was destroyed; production, 09-16, all gru1.469d)

20:44:59 ECKV_VqhS_2qGruB8u2xcg | 20:45:46 ilDJHoOvQ0GvAQkCHn5Ytg | 20:50:09 DHcjOXyPRgawimEfO8poTA | 20:54:42 veyd2JK9RMah3KHsFFmdQQ | 20:54:57 CH2SDTm5TpC_Mag-Cx5-qw

$20 Bounty

3 Replies

Railway
BOT

19 days ago

We've looked into this from our side and haven't found anything on the Railway platform that explains what you're seeing, so working it out means digging into your specific setup.

That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.

Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.

  • Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
  • Keep it private and close the thread - Nothing becomes public and the thread closes.

Status changed to Awaiting User Response Railway • 19 days ago


Railway
BOT

19 days ago

This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.

Status changed to Open Railway • 19 days ago


ozanbilal
PRO

18 days ago

Your controls make an origin crash or restart, application latency, the client link, and DNS unlikely. The client-side 1006 together with an origin socket that remains open is consistent with an intermediate proxy losing or closing the downstream leg without immediately propagating a FIN/RST upstream; the origin may only learn about it on the next I/O or keepalive. requestTimeout and headersTimeout do not govern upgraded WebSocket sockets.

On the application side, bound the impact with a heartbeat: send ws.ping() every 20–30 seconds, mark each socket alive on pong, terminate it after one missed heartbeat, and reconnect clients with jitter. Log close/error code, timestamp, x-railway-edge, x-hikari-trace, and deployment or instance context. This will not repair an edge event, but it prevents half-open sockets from lasting 120 seconds.

Because two services are hit together on the same edge or zone while HTTP and new WebSocket handshakes continue, the decisive next step is Railway-side correlation with edge drain or rotation, connection recycling, or lifetime limits for gru1 toward us-east4-eqdc4a during the listed windows. There is no application setting that can force the ingress to propagate a close frame; if the edge event is confirmed, Railway needs to fix it or route around it.


m1cdev
PRO

18 days ago

Hi! I have pretty the same issue and just opened the ticket before I noticed yours: https://station.railway.com/questions/long-lived-websocket-connections-cut-in-efad9e69

It matches what we see: the client gets 1006, the origin keeps the socket open, and our server only learns about it when its next ping goes unanswered and the edge closes our side about 15 s later.

The app side is already covered. socket.io pings every 25 s (moving to 10 s) and closes a socket that misses its pong, and clients reconnect with jitter. We log x-railway-edge and x-railway-request-id on every connection. As you say, that limits how long a half-open socket lasts, but it doesn't stop the cuts.

Ours ran 20:47–21:14 UTC the same evening, but we were observing the same issue days before.

Two unrelated services on different continents, behind at least four different edge locations, cut in the close timestamps: that points to something shared in Railway's edge rather than to either app or to one network path.

Could someone on the Railway team check that? Specifically:

  1. edge deploys, drains or rotations;
  2. connection recycling or lifetime limits;
  3. network incidents affecting several edge locations at once.

Thanks


Welcome!

Sign in to your Railway account to join the conversation.

Loading...