Production outage - network connections to production postgres db failing
adamlutz
PROOP

20 days ago

Production outage — API cannot reach Postgres over private networking

Project: warcarp · Environment: production

Services: warcarp-api (Fastify) and Postgres (Railway managed)

Started: ~15:25–15:30 UTC, 2026-09-16 · Still ongoing at: 16:47 UTC (~80 minutes)

Impact: total loss of application function. Users cannot log in, refresh tokens, or load games.

Summary

Both services report online. Postgres is healthy and idle. But **every TCP connection

from the API to postgres.railway.internal:5432 times out**, so all database queries fail

with ETIMEDOUT. Your own Postgres dashboard cannot connect either (screenshot attached),

which is why we believe this is a platform networking issue rather than an application fault.

Evidence

1. Postgres is up and healthy. Same backend process (PID 58) since at least 14:53 UTC,

routine time-triggered checkpoints continuing through 16:28:33 UTC, zero errors, no restart,

no "too many connections". Write volume collapses after ~15:28 UTC, consistent with the

database simply receiving no traffic from that point.

2. The API process is healthy. GET /health, which touches no database, returns HTTP 200

in 0.19–0.23s consistently throughout the incident. OPTIONS preflights return 204 in under

1ms. So the container, the HTTP stack and the edge are all fine.

3. Every database query times out at the socket level. From the API logs:

err={"type":"PrismaClientKnownRequestError",
     "code":"ETIMEDOUT",
     "meta":{"modelName":"RefreshToken"},
     "clientVersion":"7.8.0"}
Invalid `prisma.refreshToken.findUnique()` invocation

The same ETIMEDOUT appears for prisma.user.findUnique(), prisma.game.findMany() and

every other model. ETIMEDOUT is a TCP-level timeout, not a Prisma connection-pool timeout

(which would surface as P2024), so the client is failing to establish a socket at all.

4. Reproducible from outside. Against https://api.warcarp.com:

| Endpoint | Touches DB | Result |

|---|---|---|

| /health | no | HTTP 200 in 0.19s |

| /api/v1/games/by-code/:code | yes | no response, times out at 25–30s |

| POST /api/v1/auth/login | yes | no response, times out at 30s |

5. Your dashboard cannot connect either. The attached screenshot shows the Railway

Postgres service page: "Deployment Online" green, "Required Variables" green, and

"Database Connection — Attempting to connect to the database…" spinning indefinitely.

6. Not triggered by a deploy. The current API deployment 7c0d0e74 (commit d14cfd48)

succeeded at 14:03:16 UTC and ran normally for roughly 85 minutes before the failure began.

No deploy, config change or variable change coincides with the onset.

Connection string in use: postgresql://postgres:***@postgres.railway.internal:5432/railway

(standard private-network reference, unchanged).

Questions

  1. Why did this happen? What caused private networking between warcarp-api and

    postgres.railway.internal to stop accepting connections while both services stayed

    online and Postgres itself kept running? Was there an incident on the underlying host,

    the private network fabric, or internal DNS?

  2. Is a restart useful or pointless here? Since your own dashboard also cannot reach the

    database, we have not restarted anything, on the assumption that the fault is on the

    platform side of the private network and a restart would not help. Please confirm.

  3. How do we prevent a recurrence? Specifically:

    • Is there a recommended connection or retry configuration for Prisma on Railway private

      networking that would survive this class of interruption?

    • Should we be using the public database URL as a fallback, and is that supported?

    • What monitoring or alerting does Railway expose so we detect private-network loss

      ourselves rather than learning from user reports?

    • Is there a status or incident feed we should be subscribed to that would have shown this?

This is a live beta with real users, so any interim mitigation we can apply ourselves would

be very welcome while the root cause is investigated.

Duplicate$10 Bounty

Pinned Solution

Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.

9 Replies

Railway
BOT

20 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 20 days ago


Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.


0x5b62656e5d

Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.

adamlutz
PROOP

20 days ago

thanks trying that now, will report back


0x5b62656e5d

Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.

adamlutz
PROOP

20 days ago

thank you @0x5b62656e5d, that solved my immediate problem w/ the outage.

i would like to know the reason this happened (if it was a railway one-off) or if there's anything I can do on my side to prevent such an outage in the future. i think we were down for nearly 60 minutes unfortuantely.


adamlutz
PROOP

20 days ago

Hello, the service is down again. noticing there may be a forced migration here also, but not clear why the downtime.

Screenshot 2026-09-16 at 13.08.56.png

Attachments


adamlutz
PROOP

20 days ago

forced deployment appears to be hanging at 9 minutes?

Screenshot 2026-09-16 at 13.11.00.png

Attachments


0x5b62656e5d

Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.

adamlutz
PROOP

20 days ago

fyi that we're down again, and postres container builds are now stuck at 'queued' it appears.

i had the situation above first, where i had an auto-update that was forced (even though postrgres was healthy again), that removed the healty instance, started a migration, and effectively hung 12 minutes in w/ no healthy postgres instance left at all.


adamlutz
PROOP

20 days ago

currently it's stuck on 'migrating service' and all redeploy source image cmds are queueing up behind this it appears.


adamlutz
PROOP

20 days ago

it is stuck on the region mapping:

Screenshot 2026-09-16 at 13.26.11.png

Attachments


adamlutz
PROOP

19 days ago

@here need help urgently on this please 🙏... not clear if you're able to see these follow-ups or not. thank you.


Status changed to Duplicate brody • 19 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...