Frequent WebSocket disconnects affecting multiple independent clients
pondustech
HOBBYOP

a month ago

Hi Railway Support,

We are experiencing intermittent disconnects on long-lived WebSocket connections in our production service.

We have two GANTNER GT7 access terminals connected as WebSocket clients to the same Railway service. The terminals are physically located at different locations and use completely separate internet connections/ISPs.

We frequently see the following pattern:

WebSocket is still reported as OPEN (readyState: 1)

An application-level heartbeat is sent but receives no response

After our 5-second command timeout, the socket is still OPEN

Shortly afterwards the connection closes abnormally with WebSocket code 1006

The GT7 reconnects successfully afterwards

More importantly, we have several cases where both independent devices disconnect within seconds of each other.

Examples from August 27 (CEST):

00:55:03 – device xyz

00:55:04 – device abc

04:35:38 – device abc

04:35:42 – device xyz

I've attached a screenshot of more disconnect events.

We have also seen periods where only one connection is affected while the other continues responding normally.

We added event-loop monitoring to the Node.js process to rule out application stalls. So far, the WebSocket failures have occurred without corresponding event-loop delays, and the Node process itself remains healthy.

Is this due to Railway edge/proxy connection recycling or another networking event occurred for our service around the timestamps above?

We would particularly like to know whether long-lived WebSocket connections are expected to be periodically recycled by Railway, and whether there is a recommended configuration for production systems that depend on persistent WebSocket connectivity.

Thanks!

5580b838-76c7-4612-aef2-977463e3a53a.png

Attachments

$10 Bounty

2 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


turicamirabelamaria-art
PRO

a month ago

The evidence points strongly to Railway edge/proxy connection recycling rather than a Node.js event-loop stall or two independent client-side failures.

The key signal is the correlated disconnects: two GT7 terminals on physically separate networks/ISPs are dropping within seconds of each other, while the Node process remains healthy and no event-loop delay is recorded. Both sockets then terminate abnormally with code 1006 and reconnect successfully.

Railway has previously confirmed this behavior publicly: its edge layer can recycle long-lived connections, and WebSocket connections are not guaranteed to remain open indefinitely. When an affected edge/proxy is recycled, multiple clients can be disconnected at the same time.

So I would not try to “fix” this solely by changing the WebSocket heartbeat interval. Your application-level heartbeat is still useful for detection, but it cannot prevent an upstream edge connection from being recycled.

For production, treat WebSocket continuity as recoverable rather than permanent:

keep automatic reconnect with exponential backoff + jitter;

make sessions resumable/idempotent so the terminal can reconnect without losing application state;

detect heartbeat failure quickly and reconnect rather than waiting indefinitely for readyState to change;

log exact UTC disconnect timestamps and connection/session IDs;

if multiple independent clients drop within the same few seconds, correlate those events and escalate them to Railway as likely edge-level recycling/network events.

For this specific incident, the August 27 timestamps you supplied are exactly the kind Railway staff can use to inspect edge/proxy logs and confirm whether a proxy rotation or networking event occurred.

In short: yes, Railway edge/proxy recycling can explain this pattern. There is no application-side setting that can guarantee an indefinitely persistent WebSocket through an infrastructure recycle, so the robust design is heartbeat + fast reconnect + resumable session state.


hipsterreed
PRO

4 days ago

I had this same issue and a railway rep told me that Railway doesn't guarantee that the connection won't get dropped at the edge and that you have to build reconnect mechanism. The issue i ran into is that sometimes the edge would be unstable for up too a minute and in my usecase with users on a phone it became a terrible experience and we ended up moving that service to AWS. I would love if railway would provide better service around those long lived connections.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...