Production Postgres in crash loop: volume full ("No space left on device"), need urgent volume resize
riverajefer
PROOP

25 days ago

Hi Railway team,

Our production Postgres can't start and our app is completely down. We urgently need the volume resized, but the Live Resize option doesn't show up for us.

  • Project: High solutions (50c96440-1498-46ef-a810-e9a98dba2d2f)
  • Environment: production (a471c2c4-275b-4152-b50f-6bc4bd35d689)
  • Service: Postgres (92a7ce56-dcae-4075-9982-7c0cb5877e1a), deployment d8f41eda-8fb9-400b-a5b8-dd63134a7fbb
  • Volume: postgres-volume (052876d2-e582-4828-8199-b5ddfa3a7970), 1000 MB, mounted at /var/lib/postgresql/data
  • Plan: Pro

What happened:

At ~03:07 UTC on Sep 11 we ran VACUUM FULL on a large table right after a bulk DELETE. It ran out of disk and Postgres crashed with:

PANIC: could not write to file "pg_wal/xlogtemp.70947": No space left on device

Since then the container restarts every ~40 seconds. Crash recovery replays the WAL successfully, then fails as soon as it needs to write:

LOG: redo starts at 1/2480CC88

LOG: redo done at 1/3AFF4B80 system usage: CPU: user: 0.63 s, system: 0.98 s, elapsed: 15.83 s

FATAL: could not write to file "pg_wal/xlogtemp.68": No space left on device

LOG: startup process (PID 68) exited with exit code 1

LOG: shutting down due to startup process failure

The startup output also shows "pgbackrest: volume 878 MiB".

What we need:

Please grow the volume to 5 GB, or enable Live Resize for it. Your docs say Pro can self-serve up to 1 TB, but the option isn't shown in the volume settings.

We haven't deleted any files or restored a backup: the WAL replay completes every time, so we expect Postgres to start normally once it has free space. Please don't restore from a backup without checking with us first.

Full Postgres logs (03:14–03:21 UTC) are attached.

Thank you!

Solved

2 Replies

Railway
BOT

25 days ago

We confirmed the crash loop: WAL redo completes each cycle, then the post-recovery write to pg_wal fails with "No space left on device." Your Pro plan supports self-serve volume resize up to 1 TB, and the volume is in a READY state, so Live Resize should be available from the volume's settings page under "Volume Size." Because the volume is at 100% capacity, the resize will run offline with a brief restart, after which Postgres should complete recovery and accept connections normally. No backup restore has been performed.


Status changed to Awaiting User Response Railway • 25 days ago


Status changed to Solved riverajefer • 25 days ago


Railway

We confirmed the crash loop: WAL redo completes each cycle, then the post-recovery write to pg_wal fails with "No space left on device." Your Pro plan supports self-serve volume resize up to 1 TB, and the volume is in a READY state, so Live Resize should be available from the volume's settings page under "Volume Size." Because the volume is at 100% capacity, the resize will run offline with a brief restart, after which Postgres should complete recovery and accept connections normally. No backup restore has been performed.

riverajefer
PROOP

25 days ago

Thank you, I can resolved


Status changed to Awaiting Railway Response Railway • 25 days ago


Status changed to Solved Railway • 25 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...