2 days ago
Hi all,
We run a WebSocket service on Railway (Pro plan, us-east4, one replica, custom domain, port 8080) with clients that stay connected for hours. Several times a day, every open connection to the service is closed within the same second, while the service keeps running normally. Clients reconnect and resume within about a second, so nothing is lost, but each wave is a visible disconnect for every client. The pattern looks like this thread, which has no answer yet.
What we measured
On 2026-10-08 we ran raw WebSocket probes from two networks in Europe (a home connection in France and a server elsewhere), each through Cloudflare and directly to the Railway edge. The service sends a heartbeat and a WebSocket ping every 5 s, so connections are never idle.
All four waves hit every path within the same second:
| Time (UTC) | Client | Path | x-hikari-trace | x-railway-request-id |
|---|---|---|---|---|
| 14:49:43 | server | Cloudflare | cdg1.8vsn | |
| 14:49:43 | server | direct | ams1.cycp | kMUgR8MKS7ulo3DkCx5-qw |
| 14:49:43 | home | Cloudflare | cdg1.8vsn | |
| 15:00:23 | home | direct | cdg1.e9jw | P7pGYy1qScyQpUAjH4GxDA |
| 15:00:23 | home | Cloudflare | cdg1.e9jw | |
| 15:45:48 | server | Cloudflare | cdg1.8vsn | |
| 15:45:48 | server | direct | ams1.qkjh | YsfxnEHHR-iiBB8VGbGh5g |
| 15:45:48 | home | Cloudflare | cdg1.e9jw | |
| 15:55:59 | home | direct | cdg1.8vsn | wK-aFKjGRfClsJnsHn5Ytg |
| 15:55:59 | home | Cloudflare | cdg1.e9jw | |
- Every close is a 1006: a TCP FIN from the peer arrives 0–4 ms before it, with no WebSocket close frame.
- Frames, including the 5 s heartbeat, arrived normally until the close.
- No deploy or restart of our services at these times. The service logged nothing, and its metrics show no CPU or memory pressure. Our own deploys that day (13:55Z, 18:55Z) close with 1012 and are not in the table.
- Reconnections after a wave stay up; we don't see the 60 s re-drops of the thread above.
- The waves come in pairs about 10 minutes apart (14:49 → 15:00, 15:45 → 15:55).
Since one wave hits different edge nodes (cdg1 and ams1), two client networks, and both the Cloudflare and the direct path, we suspect a layer shared behind the edge nodes, between the edge and the us-east4 origin. We haven't confirmed that.
Questions
- Is there a periodic restart or rebalancing of the proxy layer that closes long-lived WebSockets? If so, how often, and is there a notice we can follow?
- Can such a restart drain connections gracefully (a close frame, or a delay) instead of cutting TCP?
- Would a TCP proxy avoid this layer for long-lived connections?
- Is there any setting (region, replicas, domain configuration) that reduces these closes?
Happy to share more timestamps and request ids. Thanks!
1 Replies
Status changed to Awaiting Railway Response Railway • 1 day ago
a day ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 1 day ago
2 hours ago
Your data already names the failing layer. One wave hits two edge PoPs, two networks, and both Cloudflare and direct paths in the same second, so the common piece sits behind the edge nodes toward us-east4. Your app is out: silent logs, flat metrics, and your own deploys close with 1012 while these waves close with 1006. Idle timeouts are out as well. You heartbeat every 5 s, and Railway exempts WebSockets from duration and inactivity limits. The paired waves 10 minutes apart look like a rolling restart of a shared proxy or mesh hop.
- Restart schedule or notice? Only staff can confirm from internal deploy logs. No public feed covers routine edge restarts. Ask them to correlate proxy and mesh events at 14:49, 15:00, 15:45, 15:55 UTC on 2026-10-08 using your request IDs.
- Graceful drain? No. Behavior shows FIN with no close frame, and no setting enables drain. Keep client reconnect-with-resume and file the close frame as a feature request.
- TCP proxy? Yes, best bypass. Raw TCP avoids the HTTP edge entirely. Enable it on 8080 next to your domain, CNAME a subdomain to the proxy host if wanted (Cloudflare DNS only). Costs: clients use a proxy port instead of 443, and reconnects land on random replicas.
- Other settings? EU West shortens the path but will not fix shared-layer restarts. More replicas only spread reconnects. Keep the heartbeat, add resume tokens and reconnect jitter.
Capture X-Railway-Upstream-Zone via X-Railway-Debug on upgrade. A zone change across a wave confirms origin-side rebalancing. Then run your four probes against the TCP endpoint for 48 h. Waves gone means the HTTP edge path was guilty and you have your workaround. Waves persisting means host or mesh, staff telemetry only.