a day ago
Hi Railway team,
Our Postgres service has severe intermittent disk-write latency that is not caused by our workload. It has gotten noticeably worse over the last two days, and it is now causing user-facing failures.
Setup
- Project: perpetual-embrace · Environment: production · Region: sfo
- Service: pgvector (pgvector/Postgres 18.4) with volume "pgvector-volume" mounted at /var/lib/postgresql
- Database ~1.2 GB; volume 1.6 GB / 5 GB used
- Memory use ~0.3–1.2 GB of the 8 GB limit, CPU ~0 vCPU
What we see
Postgres checkpoints that write only a handful of buffers take minutes. The worst ones are in sync/fsync. Examples from the service log (UTC):
2026-10-04 05:23:06 checkpoint complete: wrote 20 buffers … write=6.823 s, sync=36.252 s, total=79.851 s; longest=5.016 s
2026-10-04 05:33:14 checkpoint complete: wrote 19 buffers … write=26.549 s, sync=12.844 s, total=88.410 s; longest=6.449 s
2026-10-04 05:40:06 checkpoint complete: wrote 25 buffers … write=103.444 s, sync=20.553 s, total=199.227 s; longest=10.398 s
At other times the same checkpoints sync in a few milliseconds (for example 06:43 UTC: sync=0.005 s).
Share of checkpoints taking more than 5 seconds, per day (from our log):
27 Sep 28% · 28 Sep 27% · 29 Sep 33% · 30 Sep 38% · 1 Oct 26% · 2 Oct 30% · 3 Oct 50% · 4 Oct 72%
The worst checkpoint each day took 2–4.5 minutes.
It is not our load
- From 05:11 to 12:48 IST today (7.5 h, which includes the stall above), Postgres read about 32,000 blocks (~250 MB) in total and wrote about 1,400 rows. That is roughly 1 block per second, far below the 3,000 IOPS volume limit.
- On 3–4 Oct we wrote less data than on 29 Sep–2 Oct, yet the slow-checkpoint share went from ~30% to 72%.
- During the stalls, commits of tiny transactions take 5–13 s. Our app sees "Transaction already closed … commit on an expired transaction" errors, and real user uploads fail.
This looks like the degraded-storage-host / snapshot write-amplification issue described in these threads:
- https://station.railway.com/questions/severe-postgres-disk-i-o-latency-38s-c-22bb5414
- https://station.railway.com/questions/database-performance-issues-slow-stor-f6213a55
Our requests
- Please check the storage host behind pgvector-volume. If it is degraded, please migrate the volume to a healthy host. We can trigger a redeploy at a time you suggest, and a few minutes of downtime is fine.
- Are volume snapshots/backups contributing to write amplification on this host? If so, what would you recommend for our backup settings?
- Is there anything on our side (Postgres settings, service settings) you'd recommend to make this volume more resilient?
We can share more logs or run any diagnostics you need. Thanks a lot!
Aditya
NoteAlong (notealong.com)
0 Replies
Status changed to Awaiting Railway Response Railway • 1 day ago