Postgres container frozen (locked phantom state) - unreachable, redeploy failed with no logs
bureau-noir
PROOP

a month ago

Our production Postgres service became unreachable at ~23:33 UTC on 2026-08-24 and is still down. Full outage of our app (seao.groupesw.com): web and workers time out connecting to postgres.railway.internal:5432.

What we observe:

  • TCP proxy (caboose.proxy.rlwy.net:23334): TCP connects, but the PostgreSQL handshake never completes (timeout / "server closed the connection unexpectedly").
  • railway ssh --service Postgres returns "Error: Session error". The dashboard Database tab shows "Deployment Online" but "unable to connect to the database via SSH — Connection terminated unexpectedly".
  • Metrics: CPU ~0, memory decreasing (backends dying), disk 5/50 GB — no resource pressure.
  • Postgres logs from 23:33 to 23:42 UTC: "WARNING: autovacuum worker took too long to start; canceled" (x3) and "LOG: could not accept SSL connection: EOF detected" (x4). No FATAL, no OOM. These lines reached the log stream ~16 minutes late, so the host's log agent was stalled too. Nothing logged since 23:42.
  • Other services in the project (Redis, workers) and another Postgres in a different project are healthy. Nothing has been deployed on our side (last deploy 2026-08-23).

At 23:53 UTC I ran railway redeploy --service Postgres. The new deployment stayed in DEPLOYING with zero build/deploy logs and was marked FAILED by the platform at 00:26 UTC. The old deployment still shows ACTIVE/SUCCESS — the frozen container is never stopped, so the volume is never released.

Request: this matches the "locked phantom state" described in the thread "Can't Redeploy Postgre DB" (old container not shutting down during redeploy). Please evict/force-stop the frozen container so the volume is released and a new deployment can start (or reschedule the service on a healthy host). We would prefer to keep the live volume rather than restore a backup.

IDs:

  • Project: 1c83b6f8-5d67-4470-9cd6-424f23cba3df (production env fe0b2104-eac5-47dc-b43c-b8dd7d0c89c5), region us-west2
  • Service Postgres: bdd43f31-7a45-4f44-9ab1-ea772e783a2a
  • Old deployment (frozen, shows ACTIVE): 81fe49a2-8092-40f8-8c46-51ae4ea9a27b
  • Redeploy (FAILED at 00:26 UTC, no logs): f67eaa93-bdb9-4da4-9903-9f2231aada55
  • Volume instance: b6b58364-5bcc-43ea-baac-c7d18ed92fc1
Solved

2 Replies

Status changed to Awaiting Railway Response Railway • about 1 month ago


bureau-noir
PROOP

a month ago

Update: the database came back on its own at 00:46:52 UTC (same deployment restarted in place, automatic WAL recovery completed in <1 s, no data loss). If someone on your side rebooted the host, thanks.

Two questions remain if you can check:

  1. Was there a host-level event in us-west2 around 23:33 UTC that would explain the container freeze?
  2. Is there anything we should do to avoid a redeploy ending up stuck/FAILED with no logs when the old container can't be stopped?

Otherwise this can be closed.


HI there, seems to be a host issue on that machine, you or others may see some impact on that host. Working with the team to recover that workload.


Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...