a month ago
Hello Railway team,
I received a server hardware failure notification for my production PostgreSQL service in Southeast Asia (Singapore), saying recovery was in progress and the service would automatically resume. It is still unreachable more than seven hours later. The affected production database is selected in this support request.
Observed timeline (UTC):
- 2026-09-06 14:37: Normal PostgreSQL checkpoint and WAL archive success messages.
- 2026-09-06 14:45:59: Last visible database deployment log is "Mounting volume on: [path redacted]". No subsequent successful PostgreSQL startup message was visible during our check.
- 2026-09-06 21:53 (2026-09-07 06:53 KST): The backend still reports connect ETIMEDOUT to the database on port 5432, and authentication refresh requests return HTTP 503.
The Railway dashboard marks the database Online/Active, but the Database tab remains at "Attempting to connect to the database". The global status page shows Fully Operational, while the project notification reports a hardware failure affecting this service.
Could the Railway team please check the affected host and volume recovery, confirm whether PostgreSQL has restarted successfully, and provide a recovery status update or ETA? Please also advise whether we should continue waiting for automatic recovery or whether any specific customer action is needed.
This is affecting our production application's authentication and database-dependent functionality. We have not restarted, redeployed, or modified the database during this investigation.
Thank you.
1 Replies
a month ago
A hardware failure hit the host your Postgres service runs on at 14:36 UTC today. The host recovery is now complete and your volume data is intact (approximately 1.2 GB on disk), but PostgreSQL itself did not finish restarting after the container was re-placed. This occasionally happens after a host-level recovery, and the service will stay in this state until you trigger a fresh deploy.
To bring it back, open the service in your project, press Cmd+K (Ctrl+K on Windows/Linux) to open the command palette, and run "Redeploy source image". This schedules a new container on a healthy host. Redeploying does not delete, reset, or roll back your volume data, so your database contents will be there when PostgreSQL starts up and runs its WAL crash recovery.
For mission-critical workloads, we advise implementing a high-availability configuration (primary-replica setup) so a single-host failure does not take production down.
Status changed to Awaiting User Response Railway • 29 days ago
Status changed to Solved hyw950 • 29 days ago