Postgres volume fsync stalls (10–27s checkpoints) on a near-idle database — US West
henrjus
HOBBYOP

2 days ago

Project: remarkable-purpose, environment: production

Service: Postgres (ghcr.io/railwayapp-templates/postgres-ssl:18), volume: postgres-volume

Region: US West (California), 1 replica. Web service (Django/gunicorn) is in the same region, on the private network.

Since about 17:50 UTC on 2026-10-03, Postgres checkpoints that write only 4–20 buffers are taking 8–27 seconds, almost all of it waiting on fsync. Before that, the same-sized checkpoints took 0.1–2.9 s, with sync times of 0.003–0.12 s.

This began the same day as the US West Compute Outage (incident 72DDHCC1) and is still happening after that incident was marked resolved, like the other "slow volume writes in US West since Oct 3" threads posted today.

Examples from the Postgres deploy log:

Normal (2026-10-02 19:01 UTC):

checkpoint complete: wrote 4 buffers; write=0.410 s, sync=0.021 s, total=0.474 s; longest=0.012 s

Degraded:

2026-10-03 17:52 UTC wrote 13 buffers; write=1.335 s, sync=13.298 s, total=15.540 s; longest=10.906 s

2026-10-03 23:12 UTC wrote 16 buffers; write=2.915 s, sync=2.256 s, total=21.928 s

2026-10-03 23:27 UTC wrote 14 buffers; write=2.315 s, sync=5.893 s, total=25.123 s; longest=3.535 s

2026-10-03 23:37 UTC wrote 13 buffers; write=3.865 s, sync=6.899 s, total=27.088 s; longest=3.615 s

2026-10-03 23:47 UTC wrote 20 buffers; write=2.752 s, sync=5.160 s, total=9.926 s; longest=4.329 s

The database is essentially idle: CPU near 0, memory flat at ~50 MB over 7 days. So this isn't load on our side. The effect on users is that any request that commits a transaction takes from 1 to 17+ seconds (p99 on the service's Response Time chart), while requests that don't touch the database still return in ~85 ms.

This looks like degraded storage on the host behind this volume. Could you check it, or move the volume to healthy storage? Happy to schedule a short restart window if that's what it takes.

Awaiting User Response

1 Replies

Status changed to Awaiting Railway Response Railway • 1 day ago


chandrika
EMPLOYEE

2 days ago

Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE

Redeploying or restarting will not help here, we suggest holding off on both for now

Our infrastructure team is looking into it, and we'll update the incident as soon as we have more


Status changed to Awaiting User Response Railway • 1 day ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...