2 months ago
Production outage, 6 days down, 12 stores affected.
Project: imaginative-fascination / production
Service: Postgres 18.4 (postgres-ssl image)
Volume: postgres-volume, US West
Plan: Pro
Postgres has been crash-looping since Aug 12. It restarts every ~7 seconds
and never comes up:
LOG: redo starts at 82/65000028
LOG: redo done at 82/700000A8
FATAL: could not write to file "pg_wal/xlogtemp.86": No space left on device
LOG: shutting down due to startup process failure
It finishes redo and then dies writing WAL because the disk is full.
I did a Live resize to 10 GB, and the dashboard shows "Volume Size: 10.00 GB".
I also did a full redeploy after that. Still no change — every boot the
container logs:
pgbackrest: volume 433 MiB
So the volume grew but the filesystem inside didn't. Can you expand the
filesystem so Postgres can finish recovery?
One other thing while you're in there: my Postgres-PITR service shows empty,
and every boot logs "pgbackrest-watcher: iteration skipped: pg_isready=fail".
If WAL archiving was never running, that would explain pg_wal filling up.
Can you check whether archiving is actually working for this service? I don't
want to hit this again next month.
Restoring from backup onto a fresh volume is fine as a fallback if that's
faster — let me know which you'd rather do.
Mount path from the logs if it helps:
/var/lib/containers/railwayapp/bind-mounts/ef4bf085-c798-4cbb-943c-ba18402d4811/vol_rnzqhewjbxugetgo
1 Replies
2 months ago
The project has two volumes in the production environment, and the resize went to the wrong one. The volume named postgres-volume (the one you resized to 10 GB) is detached and not in use by the Postgres service. The volume actually attached to Postgres is a different one, still at 500 MB with 487 MB used, which is why every boot fails writing WAL. You can fix this by opening the Postgres service's volume settings, locating the attached volume, resizing it to 10 GB, and redeploying. Once Postgres has room to write WAL it should finish recovery. For the pgbackrest/PITR question, that is application-level Postgres configuration outside what we can diagnose here, but the community at station.railway.com can help verify your archiving setup once the database is back up.
Status changed to Awaiting User Response Railway • about 2 months ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago