a month ago
Twice in ~24 hours, every established outbound TCP connection from our
container to our managed Postgres pooler (external provider, AWS
us-west, port 6543) stopped passing data at the same instant —
mid-protocol-message — with no RST or FIN ever delivered to either end.
The flows sat half-open indefinitely. New TCP connections from the same
container kept working the whole time, and redeploying (fresh sockets)
fixed it instantly both times.
Occurrences:
- Aug 28, starting ~20:45 UTC (took two restarts to clear)
- Aug 29, 12:45 UTC — ran dead ~2h until a manual restart at 14:38 UTC
What we can show from both ends of the connection:
-
The database side was caught waiting mid-message: a backend stuck in
wait_event=ClientRead on a half-received extended-query exchange for
the entire outage. The database was otherwise idle and healthy.
-
The pooler's logs show zero errors during the stall. When we restarted
the container it logged our stuck client sockets closing ("closed
while state was idle (transaction)" ×7) and the replacement pool
connected instantly and worked.
-
The container process itself stayed healthy throughout - accepting
inbound traffic and serving responses that didn't need the database.
-
No deploy, restart, or traffic change coincided with either onset.
This looks like NAT / connection-tracking state being dropped for
established egress flows without a reset. So, questions:
-
Was there egress/NAT maintenance, migration, or failover in US West
around Aug 28 20:30–21:00 UTC or Aug 29 12:40–12:50 UTC?
-
Is there a known failure mode where long-lived established egress
flows get black-holed without an RST?
-
Can platform-side egress state for the replica above be checked at
those timestamps?
We've since added TCP keepalive, bounded connection lifetimes, and a
watchdog that detects and self-restarts, so this now self-heals - but
we'd like to understand the root cause. Full logs and pg_stat snapshots
available on request.
1 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
This is a clean, well-instrumented report — the wait_event=ClientRead capture on the DB side plus the pooler's "closed while idle" log on restart is solid evidence this isn't an app-level bug; connections just went silent mid-flow with no FIN/RST on either side, which is consistent with NAT/conntrack state being dropped rather than a normal close.
A couple of notes, with the honest caveat that confirming root cause needs Railway's internal egress/NAT telemetry, which I don't have access to:
This may be the same class of issue as another current thread — "Recurring transient private-network failures between app services and Postgres (US West)" describes almost the identical signature: established connections silently dying mid-session, DB side idle/healthy, no deploy correlated, self-recovers. That report additionally correlates one of its occurrences with a confirmed Railway incident (RL8FRJE6, "Connectivity issues in US West," Aug 10). Worth cross-referencing your two timestamps (Aug 28 ~20:45 UTC and Aug 29 ~12:45 UTC) against that thread's Jul 23 / Aug 10 / Aug 28 06:14–06:24 UTC windows — if yours line up with a broader pattern of US West egress disruptions rather than being isolated to your service, that's a much stronger signal for Railway to act on than either report alone.
Your self-heal setup (keepalive + bounded connection lifetime + watchdog) is the right mitigation regardless of root cause — TCP keepalive in particular is exactly the fix for "connection black-holed with no RST," since it forces detection of a dead path instead of waiting indefinitely.
For Railway to actually confirm the cause, they'd need to check egress/NAT/conntrack state for your specific replica at those two timestamp windows — which, combined with the other thread's timestamps, might reveal a shared underlying US-West network event.