Recurring transient private-network failures between app services and Postgres (US West) — 3rd occurrence in 5 weeks
Anonymous
PROOP

a month ago

We've had three windows in the last five weeks where our app services intermittently

lose connectivity to Postgres over private networking (postgres.railway.internal:5432),

while Postgres itself is provably healthy each time. Looking for help correlating these

against platform-side network events, since nothing appears on the status page.

Most recent window: 2026-08-28, 06:14-06:24 UTC.

Symptom: two separate backend services (independent Prisma connection pools) began

timing out simultaneously — in-flight queries hit read timeouts, and pool acquisition

timed out ("Timed out fetching a new connection from the connection pool", limit 97,

timeout 10s). Roughly 25 established connections vanished server-side during the window

(68 total vs our normal ~93). Everything self-recovered in ~10 minutes with no

intervention as dead sockets timed out.

Evidence this wasn't our workload or the database:

  • Request volume during the window was ~16% LOWER than the preceding 16 minutes

  • Postgres: zero active queries, zero blocked queries, <0.1 vCPU mid-incident,

    18 days uptime (no restart)

  • No deploys within 10 hours on any service

  • Both services affected in the same minute despite fully independent pools —

    their only shared component is the private-network path to Postgres

Earlier windows with the same signature: 2026-07-23 (bursts at 20:26, 20:36, 20:40 UTC;

Postgres logs showed "Connection reset by peer" / "SSL EOF detected" at exactly those

times while the DB was otherwise idle) and 2026-08-10 (overlapping incident RL8FRJE6).

Questions:

  1. Can Railway correlate these three windows against network events, host maintenance,

    or migrations affecting our project's placement in US West?

  2. Is there anything placement- or configuration-wise that reduces exposure to

    private-network disruptions for a latency-sensitive DB link?

  3. Established connections dying silently (no RST reaching the app) with no status-page

    signal makes this class of issue invisible until customers hit timeouts — is any

    private-network health signal planned?

Project and service IDs, exact timestamps, and connection-count data available on

request. Severity: recovered/not blocking, but the recurrence rate is becoming a

reliability concern for us.

$20 Bounty

1 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


turicamirabelamaria-art
PRO

a month ago

Your August 10 occurrence can be correlated with a confirmed Railway incident.

Railway incident RL8FRJE6 — “Connectivity issues in US West” occurred on August 10. Railway reported that elevated traffic into US West saturated network capacity, causing services to become slow or unreachable, and specifically noted that databases could be unreachable.

That makes the August 10 event very likely platform-side rather than a Prisma/Postgres workload issue.

The July 23 and August 28 windows are more interesting because I cannot find corresponding public status incidents.

Your evidence strongly argues against ordinary pool exhaustion or database overload:

two independent application services failed in the same minute;

request volume was lower, not higher;

PostgreSQL remained almost idle and never restarted;

no deploy occurred;

established DB connections disappeared;

previous occurrences produced Connection reset by peer / SSL EOF detected;

everything recovered without application or database intervention.

Railway private networking is implemented as an encrypted WireGuard mesh between services. A transient host/network/tunnel disruption therefore can kill existing TCP sessions without requiring the Postgres process itself to restart.

Railway also has precedent for localized networking failures that do not necessarily imply a region-wide outage. In a previous incident, Railway identified a single host with exhausted network capacity affecting both private and public networking.

For Railway staff, I would specifically ask them to correlate the three windows against:

physical/compute host IDs for both backend replicas and Postgres;

host migrations or maintenance;

WireGuard/private-network peer or route changes;

packet loss / interface resets;

network-capacity alarms;

workload rescheduling or placement changes;

any host-level event that did not cross the threshold for the public status page.

The timestamps to correlate are:

2026-07-23: 20:26, 20:36, 20:40 UTC

2026-08-10: overlapping Railway incident RL8FRJE6

2026-08-28: 06:14–06:24 UTC

I would not change database tuning merely to mask this. A Prisma pool limit of 97 may be worth reviewing independently, but it does not explain two independent pools losing established connections simultaneously while PostgreSQL is idle.

For the next occurrence, one useful addition would be a small continuous probe from each backend that logs:

RAILWAY_REPLICA_ID

RAILWAY_REPLICA_REGION

DNS A/AAAA results for postgres.railway.internal

raw TCP connect latency to port 5432

Prisma pool errors

timestamp in UTC

Railway exposes replica ID and region as runtime variables, so this would let you prove whether failures correlate with a specific service placement.

I don't see a documented customer-facing host-affinity/anti-affinity control that would guarantee avoidance of this class of failure. Railway does support changing regions, but moving Postgres with an attached volume requires volume migration and downtime, so I would not move regions until Railway correlates these incidents internally.

In short: August 10 is already correlated with a confirmed Railway US West networking incident. The July 23 and August 28 events have the same failure signature but need Railway's internal host/private-network telemetry to determine whether they were smaller host- or path-level events.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...