a month ago
Project: crm-cae-2
Environment: production
Service: Postgres
CRITICAL: Database deployment STUCK for 10+ minutes in CREATE_CONTAINER step.
Root cause: Postgres checkpoints taking 192+ seconds (normal = 0.3-2s),
blocking all connections. Users cannot login.
Volume: 5000 MB allocated, only 0.8 GB used (not a space issue)
Deployment status: STUCK since 12:31 UTC
Need help:
- Unblock the stuck Postgres deployment
- VACUUM ANALYZE the database
- Check checkpoint configuration
- Diagnose table bloat causing huge WAL writes
10 Replies
a month ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • about 1 month ago
a month ago
I tried aborting and redeploying as suggested, but the new deployment (b294ac9b, Postgres, production) is stuck in DEPLOYING with no logs — the infra itself is down. We got the "hardware failure" notification for the server hosting our Postgres volume. This is our production DB and our whole platform is down since ~11:57 UTC — this is urgent. Can you migrate/restore the volume or give an ETA?
Attachments
a month ago
It's an server outage, and as the notification says the team is working on it to get it fixed, just wait a bit, if its still not fixed, try Redeploy the source image
Open your postgres service > Open the command pallete (Ctrl-K) > select Redeploy source image
a month ago
I have the same problem on my service. Identical, hardware issue. 30 minutes passed, nothing happened.
a month ago
i think we need only to wait a bit more, if is a hardware issue can take more than the time stablished in an automatic message ( happens the same to me) to avoid issues with database i recomend use High availability db
a month ago
Because the node is offline Railway can't pull any data from the volume, and thus can't deploy to another mode. You can try restoring a backup from the backup tab if possible. If you have any external backups you can temporarily Create a new service with this backup. If you don't have any of these then unfortunately you're gonna have to wait.
danny
Because the node is offline Railway can't pull any data from the volume, and thus can't deploy to another mode. You can try restoring a backup from the backup tab if possible. If you have any external backups you can temporarily Create a new service with this backup. If you don't have any of these then unfortunately you're gonna have to wait.
a month ago
Hi, could you please provide an estimated time for resolution? We've already been waiting for over 70 minutes, and this outage is directly impacting our business. Our customers cannot use our service, and we're losing revenue every minute. Even an approximate ETA would help us communicate with our clients. Thank you.
danny
I'm not a Raiway staff member so i cant give an estimate, sorry
a month ago
thank you , it's resolved finnaly
Status changed to Awaiting Railway Response Railway • about 1 month ago
Status changed to Solved 0x5b62656e5d • about 1 month ago
