20 days ago
Production outage — API cannot reach Postgres over private networking
Project: warcarp · Environment: production
Services: warcarp-api (Fastify) and Postgres (Railway managed)
Started: ~15:25–15:30 UTC, 2026-09-16 · Still ongoing at: 16:47 UTC (~80 minutes)
Impact: total loss of application function. Users cannot log in, refresh tokens, or load games.
Summary
Both services report online. Postgres is healthy and idle. But **every TCP connection
from the API to postgres.railway.internal:5432 times out**, so all database queries fail
with ETIMEDOUT. Your own Postgres dashboard cannot connect either (screenshot attached),
which is why we believe this is a platform networking issue rather than an application fault.
Evidence
1. Postgres is up and healthy. Same backend process (PID 58) since at least 14:53 UTC,
routine time-triggered checkpoints continuing through 16:28:33 UTC, zero errors, no restart,
no "too many connections". Write volume collapses after ~15:28 UTC, consistent with the
database simply receiving no traffic from that point.
2. The API process is healthy. GET /health, which touches no database, returns HTTP 200
in 0.19–0.23s consistently throughout the incident. OPTIONS preflights return 204 in under
1ms. So the container, the HTTP stack and the edge are all fine.
3. Every database query times out at the socket level. From the API logs:
err={"type":"PrismaClientKnownRequestError",
"code":"ETIMEDOUT",
"meta":{"modelName":"RefreshToken"},
"clientVersion":"7.8.0"}
Invalid `prisma.refreshToken.findUnique()` invocationThe same ETIMEDOUT appears for prisma.user.findUnique(), prisma.game.findMany() and
every other model. ETIMEDOUT is a TCP-level timeout, not a Prisma connection-pool timeout
(which would surface as P2024), so the client is failing to establish a socket at all.
4. Reproducible from outside. Against https://api.warcarp.com:
| Endpoint | Touches DB | Result |
|---|---|---|
| /health | no | HTTP 200 in 0.19s |
| /api/v1/games/by-code/:code | yes | no response, times out at 25–30s |
| POST /api/v1/auth/login | yes | no response, times out at 30s |
5. Your dashboard cannot connect either. The attached screenshot shows the Railway
Postgres service page: "Deployment Online" green, "Required Variables" green, and
"Database Connection — Attempting to connect to the database…" spinning indefinitely.
6. Not triggered by a deploy. The current API deployment 7c0d0e74 (commit d14cfd48)
succeeded at 14:03:16 UTC and ran normally for roughly 85 minutes before the failure began.
No deploy, config change or variable change coincides with the onset.
Connection string in use: postgresql://postgres:***@postgres.railway.internal:5432/railway
(standard private-network reference, unchanged).
Questions
-
Why did this happen? What caused private networking between
warcarp-apiandpostgres.railway.internalto stop accepting connections while both services stayedonline and Postgres itself kept running? Was there an incident on the underlying host,
the private network fabric, or internal DNS?
-
Is a restart useful or pointless here? Since your own dashboard also cannot reach the
database, we have not restarted anything, on the assumption that the fault is on the
platform side of the private network and a restart would not help. Please confirm.
-
How do we prevent a recurrence? Specifically:
-
Is there a recommended connection or retry configuration for Prisma on Railway private
networking that would survive this class of interruption?
-
Should we be using the public database URL as a fallback, and is that supported?
-
What monitoring or alerting does Railway expose so we detect private-network loss
ourselves rather than learning from user reports?
-
Is there a status or incident feed we should be subscribed to that would have shown this?
-
This is a live beta with real users, so any interim mitigation we can apply ourselves would
be very welcome while the root cause is investigated.
Pinned Solution
20 days ago
Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.
9 Replies
20 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 20 days ago
20 days ago
Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.
0x5b62656e5d
Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.
20 days ago
thanks trying that now, will report back
0x5b62656e5d
Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.
20 days ago
thank you @0x5b62656e5d, that solved my immediate problem w/ the outage.
i would like to know the reason this happened (if it was a railway one-off) or if there's anything I can do on my side to prevent such an outage in the future. i think we were down for nearly 60 minutes unfortuantely.
20 days ago
Hello, the service is down again. noticing there may be a forced migration here also, but not clear why the downtime.
Attachments
20 days ago
forced deployment appears to be hanging at 9 minutes?
Attachments
0x5b62656e5d
Try clicking on your database, press Cmd/Ctrl + K, and select redeploy source image.
20 days ago
fyi that we're down again, and postres container builds are now stuck at 'queued' it appears.
i had the situation above first, where i had an auto-update that was forced (even though postrgres was healthy again), that removed the healty instance, started a migration, and effectively hung 12 minutes in w/ no healthy postgres instance left at all.
20 days ago
currently it's stuck on 'migrating service' and all redeploy source image cmds are queueing up behind this it appears.
19 days ago
@here need help urgently on this please 🙏... not clear if you're able to see these follow-ups or not. thank you.
Status changed to Duplicate brody • 19 days ago