a month ago
Our Postgres service (project: www.friendlypartyrental.com, production environment) is stuck in a crash loop. Deploy logs show: "invalid resource manager ID in checkpoint record" followed by "PANIC: could not locate a valid checkpoint record at 0/1FBDC5F8". Database was last known up at 2026-08-25 04:54:19 UTC. This began after several Docker image changes were applied to the Postgres service in a short window (postgres-ssl:18 -> postgres-ssl:18-latest -> postgres:18.3 -> postgres:latest -> back to postgres-ssl:18) as part of an attempted CVE-2026-15741 patch. We have no volume backups and PITR was not enabled, so we have no safe restore point. This is production data (customer orders/payments) and we cannot afford any data loss. Please advise on the safest way to recover the checkpoint/WAL state without losing committed transactions, or let us know if Railway can assist directly. Manual restarts of the deployment do not resolve it - it crashes again within seconds with the same PANIC error.
1 Replies
a month ago
Our databases are unmanaged, so automatic snapshots, WAL recovery, and PITR aren't available to restore your data unless you've enabled them on the service. Volume backups and point-in-time recovery are both native Railway features on Pro. We don't recover data lost to user-initiated actions. Going forward, enable volume backups on your stateful services so you can self-restore if this happens again.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago