Established egress TCP connections silently black-holed (no RST) — happened twice, US West
cody229
HOBBYOP

a month ago

Twice in ~24 hours, every established outbound TCP connection from our

container to our managed Postgres pooler (external provider, AWS

us-west, port 6543) stopped passing data at the same instant —

mid-protocol-message — with no RST or FIN ever delivered to either end.

The flows sat half-open indefinitely. New TCP connections from the same

container kept working the whole time, and redeploying (fresh sockets)

fixed it instantly both times.

Occurrences:

  • Aug 28, starting ~20:45 UTC (took two restarts to clear)
  • Aug 29, 12:45 UTC — ran dead ~2h until a manual restart at 14:38 UTC

What we can show from both ends of the connection:

  • The database side was caught waiting mid-message: a backend stuck in

    wait_event=ClientRead on a half-received extended-query exchange for

    the entire outage. The database was otherwise idle and healthy.

  • The pooler's logs show zero errors during the stall. When we restarted

    the container it logged our stuck client sockets closing ("closed

    while state was idle (transaction)" ×7) and the replacement pool

    connected instantly and worked.

  • The container process itself stayed healthy throughout - accepting

    inbound traffic and serving responses that didn't need the database.

  • No deploy, restart, or traffic change coincided with either onset.

This looks like NAT / connection-tracking state being dropped for

established egress flows without a reset. So, questions:

  1. Was there egress/NAT maintenance, migration, or failover in US West

    around Aug 28 20:30–21:00 UTC or Aug 29 12:40–12:50 UTC?

  2. Is there a known failure mode where long-lived established egress

    flows get black-holed without an RST?

  3. Can platform-side egress state for the replica above be checked at

    those timestamps?

We've since added TCP keepalive, bounded connection lifetimes, and a

watchdog that detects and self-restarts, so this now self-heals - but

we'd like to understand the root cause. Full logs and pg_stat snapshots

available on request.

$10 Bounty

1 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


dolapobayo42-lgtm
FREE

a month ago

This is a clean, well-instrumented report — the wait_event=ClientRead capture on the DB side plus the pooler's "closed while idle" log on restart is solid evidence this isn't an app-level bug; connections just went silent mid-flow with no FIN/RST on either side, which is consistent with NAT/conntrack state being dropped rather than a normal close.

A couple of notes, with the honest caveat that confirming root cause needs Railway's internal egress/NAT telemetry, which I don't have access to:

This may be the same class of issue as another current thread — "Recurring transient private-network failures between app services and Postgres (US West)" describes almost the identical signature: established connections silently dying mid-session, DB side idle/healthy, no deploy correlated, self-recovers. That report additionally correlates one of its occurrences with a confirmed Railway incident (RL8FRJE6, "Connectivity issues in US West," Aug 10). Worth cross-referencing your two timestamps (Aug 28 ~20:45 UTC and Aug 29 ~12:45 UTC) against that thread's Jul 23 / Aug 10 / Aug 28 06:14–06:24 UTC windows — if yours line up with a broader pattern of US West egress disruptions rather than being isolated to your service, that's a much stronger signal for Railway to act on than either report alone.

Your self-heal setup (keepalive + bounded connection lifetime + watchdog) is the right mitigation regardless of root cause — TCP keepalive in particular is exactly the fix for "connection black-holed with no RST," since it forces detection of a dead path instead of waiting indefinitely.

For Railway to actually confirm the cause, they'd need to check egress/NAT/conntrack state for your specific replica at those two timestamp windows — which, combined with the other thread's timestamps, might reveal a shared underlying US-West network event.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...