2 months ago
Summary: Our Postgres service (Postgres + PostGIS, deployed as postgis/postgis image) was found with "no active deployment" and had apparently been down since a Postgres crash-recovery log showed "database system was interrupted; last known up at [date ~3 days prior]." We've since gotten Postgres itself running again, but the database our application uses now has no tables — the schema and all data appear to be gone, even though the server is healthy.
Timeline of what we found and did:
Discovered the service had no active deployment and its volume was unmounted from it.
Found two files at the volume root: .railway-major-upgrade.lock and .railway-postgres-runtime.lock, suggesting Railway's own automated major-version-upgrade process had run on this service.
Redeployed our original image (postgis/postgis:16-3.4) — crashed on startup with: unrecognized configuration parameter "autovacuum_worker_slots" ... configuration file contains errors.
Tried postgis/postgis:17-3.4 — identical crash.
Determined autovacuum_worker_slots is a PostgreSQL 18–only parameter, meaning the on-disk data had been initialized/touched by PG18.
Used railway volume files download/upload to directly comment out that one line in postgresql.conf on the volume.
Redeployed 16 and 17 again — got a clearer error this time: FATAL: database files are incompatible with server. DETAIL: The data directory was initialized by PostgreSQL version 18, which is not compatible with this version 17.0. Confirmed the actual data pages are genuinely PG18 format.
Deployed postgis/postgis:18-3.6 (current PG18 + PostGIS build) — Postgres started successfully, completed a normal WAL crash-recovery ("redo done," "checkpoint complete"), and reported "database system is ready to accept connections."
However: connecting to the database our app uses shows zero tables — no public schema objects at all, just the empty defaults. No other database or schema on the instance contains our data either.
Looking at the raw data directory (/pgdata/base/), there are 4 numbered subdirectories. One of them has a last-modified timestamp matching the last-known-good time Postgres itself reported during recovery, while the others show today's date from our recovery attempts — possibly leftover/orphaned data from before whatever upgrade process ran. We stopped short of touching raw files further to avoid complicating any recovery.
What we're hoping someone can help with:
Whether the 16→18 major-version upgrade that appears to have run on this service completed correctly, or failed partway through.
Whether there's a recoverable pre-upgrade data directory, snapshot, or internal backup from before the outage, even though the dashboard's Backups tab shows empty for us.
Whether anyone else has hit .railway-major-upgrade.lock leaving a service in this state, and how they resolved it.
Happy to provide more diagnostic detail as needed.
2 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
We dug into this at the platform level, and here's what actually happened. Railway's automated major-version upgrade never ran on your service. It only operates on Railway's own Postgres images, so your postgis/postgis deployment isn't eligible, and there's no upgrade record anywhere for this project. The .railway-major-upgrade.lock and .railway-postgres-runtime.lock files you found are written by the official Postgres template image on every boot. They're startup markers, not evidence of an upgrade.
The volume attached to your postgis service today was created by one of your Postgres template deploys on Aug 15 at 20:46 UTC. What's on it is that template's first boot: a brand new, empty PG18 cluster with only the default databases. That's why 16 and 17 refuse to start on it, and why there are no application tables. It was never your data volume.
Your data was on a separate 5GB volume that was deleted from your account on Aug 15 at 20:59 UTC, six minutes after the postgis service was created. Deletions get a 48-hour grace window and an email to the workspace admins, then the data is permanently removed. That happened on Aug 17, so I'm sorry, but there's no copy left on our side to restore.
The deletion came from your own account, in the middle of a quick sequence of service creates and deletes that evening. If you're running Claude Code or another agent against this account, its history from Aug 15 around 20:45 to 21:15 UTC will show the exact sequence. Enabling volume backups on your database going forward would make something like this recoverable.
Status changed to Awaiting User Response Railway • about 2 months ago
Status changed to Solved dizzydes90 • about 2 months ago
2 months ago
thank you for the quick response!
Status changed to Awaiting Railway Response Railway • about 2 months ago
Status changed to Solved Railway • about 2 months ago