Postgres volume fsync ~1 second at under 1 MB/s writes (production impact now)
treybenfunderburg
PROOP
an hour ago
Hi Railway team,
Our production Postgres volume is flushing to disk extremely slowly right now, and it's causing timeouts across our app. We need help urgently.
- Postgres service: 4e955526
- Volume: postgres-volume 658dc20e
- Postgres 17.11, container memory 24 GB
What we measure from inside Postgres:
- Right now (Oct 3, ~21:23 UTC) each WAL fsync takes ~1,046 ms (25 fsyncs in 30 s). Earlier today it was 562–672 ms. Our 8-day average is 11 ms; healthy is 1–2 ms.
- Our write load is tiny at the same time: WAL ~0.38 MB/s, background writer ~0.7 MB/s, checkpoint writes 0/s. The disk is flushing ~100x slower than normal with almost nothing to flush.
- Effect: every commit waits behind a ~1 s flush, 10–30 sessions wait on WAL at once, and our app requests time out.
- Started around 20:13 UTC today and is ongoing. Earlier stall windows today: 09:00–09:30 and 13:15–13:50 UTC.
- This looks like the storage-layer issue in thread "Severe storage I/O degradation on Postgres host" from Aug 29.
Questions:
- Is there a storage incident or a noisy neighbor on the host behind this volume?
- Are we being throttled, and what IOPS/throughput is this volume provisioned for?
- Can you fix it now (move the volume to a healthy host or a faster tier), and what would it cost?
We have not restarted the database and will not unless you advise it.
Thanks,
Trey, Carwash Blitz
0 Replies
Status changed to Awaiting Railway Response Railway • about 1 hour ago