18 days ago
Service: api, region EU West (Amsterdam), 1 replica, custom domain api2.citizenpay.xyz -> port 8080, WebSockets over HTTP/1.1 at GET /v2/terminals/ws.
Our service holds about 19 long-lived WebSocket connections from ESP32 devices. Several times a day every one of those connections is closed at the same time from the client side, while the connection between your proxy and our container stays open. We have ruled out our container and our code; the evidence points at the edge.
What we observe
-
Waves of four. Each incident is four fleet-wide drops spaced about 8.5, 8.0 and 7.5 minutes apart, then quiet for one to nine hours. Since 2026-09-11 we count about 25 incidents. Today (2026-09-17, UTC): 13:46:53, 13:55:38, 14:03:43, 14:11:21. Earlier today: 02:56:22, 03:04:53, 03:12:53, 03:20:23 and 04:38:45, 04:47:14, 04:55:23, 05:02:54.
-
The client sees a FIN. Terminals log a clean TCP close (ESP_ERR_ESP_TLS_TCP_CLOSED_FIN) and reconnect within about one second. Reconnects succeed immediately, so the container was up and reachable throughout.
-
The upstream leg is not closed. In our application log for 13:46:53 and 13:55:38 there are zero read errors or EOFs on the old connections and zero ping timeouts (we ping every 30 s with a 10 s pong timeout). Every drop appears only as "replaced stale connection": the client arrived on a new connection while our side of the old one was still open and healthy. Our process did not restart; the one deploy today was at 13:56:23, after the second wave had begun.
-
Replacement connections are cut again. Within each wave a terminal is cut two or three times over about 90 seconds. Connections open for 8h43m, 5h17m, 3h23m and 2h36m were cut at 13:46-13:47; their replacements were cut again after 17 s, 18 s, 32 s, 40 s, 47 s and, notably, four of them at exactly 60.0-60.1 s. That matches the documented 60-second idle timeout for HTTP/1.1 connections, which the docs say does not apply to WebSockets. It looks as if, during these windows, upgraded connections are treated as idle HTTP/1.1 keep-alives.
-
Not resource related. Container CPU is flat near zero and memory is under 100 MB all week (limits 32 vCPU / 32 GB). No Serverless, no CDN caching, no edge rules, Under Attack mode off.
Questions
- Is the edge rolling proxy instances, or reloading configuration, on a cadence that would produce four events about 8 minutes apart? The pattern is too regular to be network noise.
- Why do upgraded WebSocket connections lose their idle-timeout exemption in these windows?
- Is this service on the legacy edge or on Railway Metal, and would moving change the behaviour?
- If the edge cannot keep these connections stable, is the TCP proxy the recommended path for long-lived device connections, and does it have the same behaviour?
We can provide the application log excerpts and per-connection lifetimes for any of the timestamps above.
1 Replies
18 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 18 days ago
18 days ago
it makes sense that the connections are dropped, but the reconnections being dropped so soon after sound more like a client side issue. are you sure the esp32 code is correctly handling the reconnects?