2 months ago
Postgres service stuck — instance INITIALIZING, will not start. Production down 70+ min.
Project: 4b65e23d-cd9b-40d2-b101-332513109b38 (Domestique / production)
Service: 35879c54-1512-472c-bf5e-ad2e287eb81d (Postgres)
Volume: 378c2d73-1f18-46de-b8d9-9335c4878224 (postgres-volume)
Image: ghcr.io/railwayapp-templates/postgres-ssl:18
Region: europe-west4-drams3a
Timeline (UTC, 2026-08-18):
16:41:55 Postgres wrote a normal completed checkpoint — then went silent.
No shutdown message, no error, no PANIC. Zero log output after this.16:49 Every query from the backend service began failing at the connection
layer. Postgres logged NOTHING: no "too many clients", no disk error.
The queries never reached it.
From outside, TCP to caboose.proxy.rlwy.net:38825 is accepted, but a
raw Postgres SSLRequest (00 00 00 08 04 D2 16 2F) gets no reply.
The dashboard still reported the service "Online" throughout.17:28+ Four redeploys. All instances stuck in INITIALIZING, zero deploy logs:
33a9a94a-e43b-4cb7-b26d-5930fe88329e REMOVED
b9ad6941-0693-437d-86cd-8d3cd37738e2 REMOVED
1b442523-6e3f-408a-9b91-83a740876a45 REMOVED
7b15f9db-aaa7-4227-bf84-91a3d7254502 INITIALIZING (still)Ruled out on our side: no healthcheckPath, image digest confirmed pullable from
ghcr.io (HTTP 200), volume reports "Ready" and attached, under 5 GB.
Two questions:
-
Please get the instance started — what is blocking it?
-
The volume reports 0MB used while the service is down. Please confirm the
data on postgres-volume is intact. There are no backups on this volume,
so this is the only copy.
2 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
This is an incident on our side. One of our servers hit a hardware fault and went offline right after that last checkpoint, which is why Postgres went silent and the redeploys are stuck in initializing. On your volume: your data is intact. The 0 MB reading is just because nothing is mounted to it right now, not data loss, and the volume persists through this. We're moving it onto a healthy machine to bring it back up, and we're actively on it.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
The reattach is fully complete and it is safe to resume use now. No data has been affected. There is unfortunately a small risk of 2-5 mins downtime in the next 24h as we drain the stacker completely to prevent recurrence. My apologies for this.
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago