Postgres unresponsive since 16:42 UTC, redeploys stuck/failing (EU West) — production down
4roomkz-debug
HOBBYOP

2 months ago

Our production Postgres has been down for ~1.5 hours. I cannot recover it

from the dashboard or the CLI.

Project: positive-curiosity

Environment: production

Service: Postgres (EU West)

Project ID: 54838b29-47c9-4c3d-a37b-b65f5f1b987f

Service ID: 0958206c-e13a-4ca5-a492-86ef476c2d0b

Timeline (UTC, 2026-08-18):

16:42 Postgres stopped emitting logs entirely. It had been writing

   checkpoints every 5 minutes all day; the last log line is 16:42:08.

   The service still reported status SUCCESS.

16:42+ All connections fail. Timeouts over the private network

   (postgres.railway.internal:5432) and ECONNRESET over the public proxy

   (shinkansen.proxy.rlwy.net:14636). TCP connects, but the Postgres

   handshake never completes.

16:42 Our app service began crash-looping on DB connect: 11 restarts,

   now CRASHED (deployment a0369aad-578d-4a07-b998-2fe0c0424cf8).

   The dashboard shows "Problem processing request" on the Postgres

   service page, so the UI cannot act on this service either.

17:18 I triggered a redeploy: 00deeac6-6cb0-4949-aca9-2b5b25c5b8f0.

   It stayed in DEPLOYING with zero logs and ended FAILED.

17:58 Another deployment started: 62dff165-d6af-4f55-a413-21dafdec1ac9.

   Still DEPLOYING, still no logs.

It looks like the new container cannot start — possibly the previous wedged

instance is still holding the volume. Please recover the service.

Data integrity is the priority. We have deleted nothing and will not run

railway down on the database.

This is a production outage affecting a live corporate training programme

with 60 active users.

Solved

1 Replies

Status changed to Awaiting Railway Response Railway • about 2 months ago


dizzydes90
EMPLOYEE

2 months ago

The server hosting your Postgres service went offline at 17:14 UTC due to a hardware event, and the workload was migrated to a healthy server. The first redeploy at 17:18 UTC failed because the old server was still unreachable, and the second at 17:58 UTC was delayed in starting its container, but it has now completed successfully. Postgres performed automatic crash recovery and is accepting connections again. Your volume data is intact (423 MB, READY state).

Your app service is still on its pre-outage deployment and crash-looping on database connection timeouts. Redeploy it (Cmd/Ctrl+K, then select the latest commit) and it should reconnect to Postgres normally.


Status changed to Awaiting User Response Railway • about 2 months ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...