postgres-volume reports 5000 MB but the ext4 filesystem on it is only 434 MiB
penbrookecapital
PROOP

a month ago

postgres-volume reports 5000 MB but its ext4 filesystem is only 434 MiB. The block device was expanded; the filesystem never was. Postgres filled the 434 MiB it can actually see and crash-looped, taking our production database down 2026-09-09 12:50-15:14 UTC.

This is the same bug your team has already fixed four times. Please apply the same remedy: host-side resize2fs on the volume, then redeploy the Postgres service.

  • postgre-sql-crash-loop-after-disk-full-82a101ad - live resize 500 MB -> 5 GB stuck at 500 MB. angelo-railway expanded the filesystem and redeployed; data intact.
  • postgres-volume-resized-to-5-gb-but-reco-e183c31c - sam-a: "the volume was resized to 5 GB at the ZFS level, but the ext4 filesystem inside it was never expanded to match." Staff expanded it; data preserved.
  • volume-filesystem-not-resized-to-match-b-4587abcc - 40 GB volume, 434 MB ext4. sam-a acknowledged the same mismatch.
  • pro-same-ext4-live-resize-bug-as-solve-70d67be4 - brody: "We've expanded it manually and redeployed your Postgres service." Fixed ~5 min after the report; data intact.

Ours is the same 500 MB -> 5 GB transition as 82a101ad and e183c31c, and our filesystem is 434 MB - the same figure 4587abcc reported on a 40 GB volume, so this looks like a common initial size whose resize step silently does not run.

Evidence (from inside the container)

Block device is 5 GB:

$ cat /sys/class/block/zd7312/size

9773056        # x 512 B = 5,003,804,672 B = 5.00 GB

Filesystem on it is 434 MiB:

$ df -B1 /var/lib/postgresql/data

/dev/zd7312  454299648  444121088  114688  100%  /var/lib/postgresql/data

$ mount | grep postgresql

/dev/zd7312 on /var/lib/postgresql/data type ext4 (rw,relatime,discard,stripe=8192)

454,299,648 B of filesystem on a 5,003,804,672 B device - 9% of the device, a 10.9x shortfall.

Postgres then crash-looped. WAL replay completed every cycle; it died only on the end-of-recovery checkpoint, unable to write an 8-byte file:

PANIC: could not write to file "pg_logical/replorigin_checkpoint.tmp": No space left on device

LOG:  checkpointer process (PID 81199) was terminated by signal 6: Aborted

LOG:  all server processes terminated; reinitializing

Railway dashboard and API don't show this. railway volume list reported "Storage used: 499MB/5000MB" throughout, and volumeInstance shows sizeMB 5000 / currentSizeMB 499 / state READY - they measure the zvol allocation, not the filesystem. We had no warning.

A redeploy did not resize it. df read 434M 424M 0 100% both before and after railway redeploy --service Postgres. The container had been up 19 days (since 2026-08-21 12:34 UTC) before that, so a container-start resize did not run either.

We cannot fix it ourselves. No e2fsprogs in the image; no device node inside the container (/dev/zd7312 absent) and no CAP_MKNOD to create one (mknod: Operation not permitted); no CAP_SYS_RESOURCE, so EXT4_IOC_RESIZE_FS against the mount point fails too (CapEff: 00000000800405fb). Neither railway volume update nor the GraphQL volumeInstanceUpdate mutation accepts a size. This matches sam-a on e183c31c: it "requires platform-level filesystem expansion that users cannot perform independently."

Only this volume is affected. web-volume is correctly 4.5 GiB on its 5 GB allocation and redis-volume 433 MiB on its 500 MB, so the resize normally works.

Current state: we freed space by moving two recycled WAL segments off the volume and truncating a telemetry table. Running at 434M / 323M used / 102M avail. This is a stopgap - the filesystem is still 434 MiB and will fill again.

What we need

  1. Expand the ext4 filesystem on our Postgres volume instance to match its 5 GB block device, then redeploy the service - as in 82a101ad, e183c31c and 70d67be4.
  2. Confirm what should trigger that resize and why it did not run, since four public reports suggest it can fail silently on upgrade.
  3. Consider surfacing filesystem-level usage in the dashboard/API, or alerting when it diverges from the zvol. The metric read 10% while the database was 100% full and down.

Project novanet / environment production / service Postgres / mount /var/lib/postgresql/data. Volume and volume-instance IDs available on request.

Solved

1 Replies

Status changed to Awaiting Railway Response Railway • 27 days ago


sam-a
EMPLOYEE

a month ago

The filesystem on your Postgres volume has been expanded to match the 5 GB volume. This was done live, so no redeploy or restart was needed. df inside the container should now show roughly 4.4 GB total with about 4.1 GB free, and the space you freed as a stopgap can be used again.

On why the resize did not apply: the original resize is outside the history we retain, so I cannot tell you with certainty what failed for this volume. The other volumes in your project resized correctly, which matches what you observed.

Sorry for the outage this caused.


Status changed to Awaiting User Response Railway • 26 days ago


Railway
BOT

19 days ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • 19 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...