a month ago
Our production Postgres service became unreachable at ~23:33 UTC on 2026-08-24 and is still down. Full outage of our app (seao.groupesw.com): web and workers time out connecting to postgres.railway.internal:5432.
What we observe:
- TCP proxy (caboose.proxy.rlwy.net:23334): TCP connects, but the PostgreSQL handshake never completes (timeout / "server closed the connection unexpectedly").
railway ssh --service Postgresreturns "Error: Session error". The dashboard Database tab shows "Deployment Online" but "unable to connect to the database via SSH — Connection terminated unexpectedly".- Metrics: CPU ~0, memory decreasing (backends dying), disk 5/50 GB — no resource pressure.
- Postgres logs from 23:33 to 23:42 UTC: "WARNING: autovacuum worker took too long to start; canceled" (x3) and "LOG: could not accept SSL connection: EOF detected" (x4). No FATAL, no OOM. These lines reached the log stream ~16 minutes late, so the host's log agent was stalled too. Nothing logged since 23:42.
- Other services in the project (Redis, workers) and another Postgres in a different project are healthy. Nothing has been deployed on our side (last deploy 2026-08-23).
At 23:53 UTC I ran railway redeploy --service Postgres. The new deployment stayed in DEPLOYING with zero build/deploy logs and was marked FAILED by the platform at 00:26 UTC. The old deployment still shows ACTIVE/SUCCESS — the frozen container is never stopped, so the volume is never released.
Request: this matches the "locked phantom state" described in the thread "Can't Redeploy Postgre DB" (old container not shutting down during redeploy). Please evict/force-stop the frozen container so the volume is released and a new deployment can start (or reschedule the service on a healthy host). We would prefer to keep the live volume rather than restore a backup.
IDs:
- Project: 1c83b6f8-5d67-4470-9cd6-424f23cba3df (production env fe0b2104-eac5-47dc-b43c-b8dd7d0c89c5), region us-west2
- Service Postgres: bdd43f31-7a45-4f44-9ab1-ea772e783a2a
- Old deployment (frozen, shows ACTIVE): 81fe49a2-8092-40f8-8c46-51ae4ea9a27b
- Redeploy (FAILED at 00:26 UTC, no logs): f67eaa93-bdb9-4da4-9903-9f2231aada55
- Volume instance: b6b58364-5bcc-43ea-baac-c7d18ed92fc1
2 Replies
Status changed to Awaiting Railway Response Railway • about 1 month ago
a month ago
Update: the database came back on its own at 00:46:52 UTC (same deployment restarted in place, automatic WAL recovery completed in <1 s, no data loss). If someone on your side rebooted the host, thanks.
Two questions remain if you can check:
- Was there a host-level event in us-west2 around 23:33 UTC that would explain the container freeze?
- Is there anything we should do to avoid a redeploy ending up stuck/FAILED with no logs when the old container can't be stopped?
Otherwise this can be closed.
a month ago
HI there, seems to be a host issue on that machine, you or others may see some impact on that host. Working with the team to recover that workload.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago