Postgres crash-looping: "No space left on device" in pg_wal after volume resize
matt13342
PROOP

2 months ago

Production outage, 6 days down, 12 stores affected.

Project: imaginative-fascination / production

Service: Postgres 18.4 (postgres-ssl image)

Volume: postgres-volume, US West

Plan: Pro

Postgres has been crash-looping since Aug 12. It restarts every ~7 seconds

and never comes up:

LOG: redo starts at 82/65000028

LOG: redo done at 82/700000A8

FATAL: could not write to file "pg_wal/xlogtemp.86": No space left on device

LOG: shutting down due to startup process failure

It finishes redo and then dies writing WAL because the disk is full.

I did a Live resize to 10 GB, and the dashboard shows "Volume Size: 10.00 GB".

I also did a full redeploy after that. Still no change — every boot the

container logs:

pgbackrest: volume 433 MiB

So the volume grew but the filesystem inside didn't. Can you expand the

filesystem so Postgres can finish recovery?

One other thing while you're in there: my Postgres-PITR service shows empty,

and every boot logs "pgbackrest-watcher: iteration skipped: pg_isready=fail".

If WAL archiving was never running, that would explain pg_wal filling up.

Can you check whether archiving is actually working for this service? I don't

want to hit this again next month.

Restoring from backup onto a fresh volume is fine as a fallback if that's

faster — let me know which you'd rather do.

Mount path from the logs if it helps:

/var/lib/containers/railwayapp/bind-mounts/ef4bf085-c798-4cbb-943c-ba18402d4811/vol_rnzqhewjbxugetgo

Solved

1 Replies

Railway
BOT

2 months ago

The project has two volumes in the production environment, and the resize went to the wrong one. The volume named postgres-volume (the one you resized to 10 GB) is detached and not in use by the Postgres service. The volume actually attached to Postgres is a different one, still at 500 MB with 487 MB used, which is why every boot fails writing WAL. You can fix this by opening the Postgres service's volume settings, locating the attached volume, resizing it to 10 GB, and redeploying. Once Postgres has room to write WAL it should finish recovery. For the pgbackrest/PITR question, that is application-level Postgres configuration outside what we can diagnose here, but the community at station.railway.com can help verify your archiving setup once the database is back up.


Status changed to Awaiting User Response Railway • about 2 months ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...