Postgres volume I/O stalls — tiny checkpoints taking 1–4 minutes, getting worse (project perpetual-embrace, sfo)
adityanotealong
HOBBYOP

a day ago

Hi Railway team,

Our Postgres service has severe intermittent disk-write latency that is not caused by our workload. It has gotten noticeably worse over the last two days, and it is now causing user-facing failures.

Setup

  • Project: perpetual-embrace · Environment: production · Region: sfo
  • Service: pgvector (pgvector/Postgres 18.4) with volume "pgvector-volume" mounted at /var/lib/postgresql
  • Database ~1.2 GB; volume 1.6 GB / 5 GB used
  • Memory use ~0.3–1.2 GB of the 8 GB limit, CPU ~0 vCPU

What we see

Postgres checkpoints that write only a handful of buffers take minutes. The worst ones are in sync/fsync. Examples from the service log (UTC):

2026-10-04 05:23:06 checkpoint complete: wrote 20 buffers … write=6.823 s, sync=36.252 s, total=79.851 s; longest=5.016 s

2026-10-04 05:33:14 checkpoint complete: wrote 19 buffers … write=26.549 s, sync=12.844 s, total=88.410 s; longest=6.449 s

2026-10-04 05:40:06 checkpoint complete: wrote 25 buffers … write=103.444 s, sync=20.553 s, total=199.227 s; longest=10.398 s

At other times the same checkpoints sync in a few milliseconds (for example 06:43 UTC: sync=0.005 s).

Share of checkpoints taking more than 5 seconds, per day (from our log):

27 Sep 28% · 28 Sep 27% · 29 Sep 33% · 30 Sep 38% · 1 Oct 26% · 2 Oct 30% · 3 Oct 50% · 4 Oct 72%

The worst checkpoint each day took 2–4.5 minutes.

It is not our load

  • From 05:11 to 12:48 IST today (7.5 h, which includes the stall above), Postgres read about 32,000 blocks (~250 MB) in total and wrote about 1,400 rows. That is roughly 1 block per second, far below the 3,000 IOPS volume limit.
  • On 3–4 Oct we wrote less data than on 29 Sep–2 Oct, yet the slow-checkpoint share went from ~30% to 72%.
  • During the stalls, commits of tiny transactions take 5–13 s. Our app sees "Transaction already closed … commit on an expired transaction" errors, and real user uploads fail.

This looks like the degraded-storage-host / snapshot write-amplification issue described in these threads:

Our requests

  1. Please check the storage host behind pgvector-volume. If it is degraded, please migrate the volume to a healthy host. We can trigger a redeploy at a time you suggest, and a few minutes of downtime is fine.
  2. Are volume snapshots/backups contributing to write amplification on this host? If so, what would you recommend for our backup settings?
  3. Is there anything on our side (Postgres settings, service settings) you'd recommend to make this volume more resilient?

We can share more logs or run any diagnostics you need. Thanks a lot!

Aditya

NoteAlong (notealong.com)

Awaiting Railway Response

0 Replies

Status changed to Awaiting Railway Response Railway • 1 day ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...