Postgres crash loop — catatonit: failed to exec pid1 — production down
egloz93
HOBBYOP

a month ago

Hi Railway team — my production Postgres service is in a crash loop and I need help repairing the container while preserving my data

volume.

Service ID: aa3416c4

Project URL: (https://railway.com/project/86c0434e-a6b8-4341-8c22-8d0de30024b3/service/e3aaabfe-c136-45e5-8cca-c62616faba00/database?environmentId=39499d44-4a3f-4ea6-a40f-216da9b1d5d4)

Impacted service: Postgres

Symptom:

The Postgres deployment shows "Online" but no connection can be established:

  • My Node app: Error: Connection terminated due to connection timeout
  • Railway's own data browser: "Unable to connect to the database via SSH — Connection terminated unexpectedly"

Deploy logs (repeating every few seconds):

Mounting volume on: /var/lib/containers/railwayapp/bind-mounts/5c10aa98-7d43-4d40-be5a-d9b31d0c56b8/vol_6pnlfyz26ur6b9lu

ERROR (catatonit:2): failed to exec pid1: No such file or directory

The volume mounts successfully — my database files are intact on vol_6pnlfyz26ur6b9lu. The problem is the Postgres binary isn't

executing at PID 1. Container is stuck in restart loop.

Timeline:

  • Working normally through Friday Sept 5
  • Sometime over the weekend, writes stopped and the container went into this crash loop
  • No deployments were made on our side in this window

Ask:

Please rebuild the Postgres container against the existing volume vol_6pnlfyz26ur6b9lu without touching the data. Do NOT restore from

snapshot or reset the volume unless the existing one is confirmed unrecoverable.

Impact:

Production admin portal is degraded (users can log in but can't see any data). This is customer-facing. Please prioritize.

Happy to give you shell access, run diagnostics, or migrate to a fresh instance if needed — as long as the data volume is preserved or

migrated intact.

Thanks — Eddie

Solved

1 Replies

Railway
BOT

a month ago

Your volume is intact (333 MB used, state READY), so your database files are safe. To resolve the catatonit: failed to exec pid1 error, open the Postgres service, press Cmd+K (or Ctrl+K) to bring up the command palette, and choose "Redeploy source image" to re-pull a fresh container image. A regular redeploy reuses the cached image and will not fix this.


Status changed to Awaiting User Response Railway • 27 days ago


Railway
BOT

20 days ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • 20 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...