Postgres stuck "exited" after hardware-failure notice marked resolved (7 days down, no backups)
girlssailing
HOBBYOP

2 months ago

Project: GirlsSailing Ops, environment: production, service: Postgres-hCNm.

Issue: On 2026-08-10 I received an in-app notification "Server Outage Affecting Your Service" saying a server running Postgres-hCNm had a hardware failure and Railway was working to restore it. That notification is now shown as Resolved (7d ago), but the service has not come back:

  • The service's Console tab shows: "Your service's container is not running (status: exited). Deploy or restart your service, then try again."
    • I manually triggered Restart on the deployment (b2698899) — it did not come back up; Console still reports status: exited.
    • Deploy Logs for this deployment show "No logs in this time range" even immediately after the restart attempt, so I can't see a crash reason.
    • The volume (postgres-volume-eM4X) still shows ~400MB used, so the data does not appear to have been wiped.
    • Our production app (GirlsSailingOps service) is throwing PrismaClientKnownRequestError / ETIMEDOUT trying to reach this database, so the app is fully down.
    • We're on the Hobby plan, so there are no automatic backups to restore from — this is our only copy of the data.

Given it's been 7 days since the outage was marked resolved with no actual recovery, could someone take a look at the underlying host/volume for this Postgres service and help get it back online, or confirm whether the data is recoverable? Happy to provide any additional logs/IDs needed.

Solved

2 Replies

Railway
BOT

2 months ago

Your Postgres service was affected by a server hardware failure on 2026-08-10 and was migrated to a healthy host, but the container image was left in a stale state, which is why it exited and a normal restart did not bring it back. Your volume data is intact (the volume is attached, in READY state, and holding data). To fix this, open the Postgres-hCNm service, press Cmd+K (or Ctrl+K) to open the command palette, and select "Redeploy source image", which re-pulls a fresh container image. A normal redeploy or restart will not work here because it reuses the same stale image.


Status changed to Awaiting User Response Railway • about 2 months ago


girlssailing
HOBBYOP

2 months ago

Confirmed fixed — "Redeploy source image" from the command palette brought the container back up. Verified via the Console (live shell connected) and by querying a table directly (row counts match what I'd expect, data fully intact). App is back online too. Thanks for the fast diagnosis!


Status changed to Awaiting Railway Response Railway • about 2 months ago


Status changed to Solved Railway • about 2 months ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...