Recurring extreme COMMIT/fsync latency on Postgres (single-instance, volume-backed) — up to 43s, worsening trend
feelnopain
HOBBYOP

2 months ago

We're seeing recurring, severe stalls on plain COMMIT (and occasionally other statements — INSERT, DELETE, SELECT FOR UPDATE) with no application-level cause. Confirmed via log_min_duration_statement=200 and railway logs -s Postgres --json:

  • 2026-07-12: bare COMMIT, 20016ms, zero concurrent traffic on the DB at the time (ruled out lock contention, replication, WAL archiving — nothing else was connected).
  • 2026-07-13: 07:29:16 UTC bare COMMIT, 20016ms; 10:25:52 UTC an activity_log INSERT (10928ms) and its COMMIT (10157ms) stalled back-to-back on the same connection.
  • 2026-07-20: DELETE on cart_items (7.4s) and a bare COMMIT (4.6s).
  • 2026-07-21: INSERT on lockers, 16.5s.
  • 2026-07-22: INSERT on activity_log, 12.5s.
  • 2026-07-25 08:43 UTC: COMMIT (9.7s) blocking a lock on an orders row (8.6s downstream wait).
  • 2026-07-25 17:04 UTC: bare COMMIT stalled 28.8s. Because a lockForUpdate() on a carts row was held across that COMMIT, three subsequent add-to-cart requests for the same cart queued behind it (23.6s, 11.2s, 4.3s waits) — this is what surfaced

app-side as "add to cart very slow," with PHP-FPM's slow-log catching 5-6s+ requests and the frontend's retry logic turning it into duplicate POSTs / 499s.

  • 2026-07-26 07:22:07 UTC: bare COMMIT, 42996.719ms (~43s) — new record.
  • 2026-07-26 07:34:23 UTC: bare COMMIT, 22971.856ms, same morning, same cascading lock-wait pattern on the carts table (wait chains up to 32.9s).

We checked and ruled out:

  • Checkpoint I/O as a trigger (checkpoints logged just before/after each incident were themselves fast, 2-4s, and didn't overlap the stall windows).
  • Application-level causes (no long-held locks from our own code paths preceding these; the holder transaction's own COMMIT is what's stalling, not something the app is doing wrong).

With synchronous_commit=on and no replica/archiving to sync against, the only thing a bare COMMIT waits on is the local WAL fsync — so this points to storage/disk-layer latency on the volume backing this Postgres instance, not anything at the

SQL/application layer.

We also found what looks like the same failure mode reported by a different Railway user on a different project: https://station.railway.com/questions/severe-postgres-disk-i-o-latency-38s-c-22bb5414

Given the trend (20s → 28.8s → 43s over two weeks) and the external report suggesting a shared-storage-host issue rather than something specific to our instance, could you:

  1. Check whether our Postgres volume is on a degraded storage host (similar to the linked report), and migrate it if so.
  2. Let us know if there's a way to get ahead of this proactively (host health monitoring, a dedicated/non-shared volume tier, etc.) rather than reacting to individual stalls.

Happy to provide our project/service IDs, exact log excerpts, or a specific timestamp window if that helps you correlate against host-level metrics on your end.

$10 Bounty

1 Replies

Railway
BOT

2 months ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 2 months ago


Anonymous
FREE

2 months ago

Capture pg_stat_activity during a stall and confirm the backend is waiting on WalSync.

Compare pg_stat_wal wal sync time before and after to prove durable storage latency.

Checkpoint tuning cannot fix a volume that takes tens of seconds to fsync.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...