: Postgres volume fsync 10–30 s in US West since Oct 3 (after incident 72DDHCC1), production POS affected
Anonymous
HOBBYOP

2 days ago

Our production Postgres has had severe disk sync latency since ~12:00 UTC on Oct 3, starting with incident 72DDHCC1, and it got worse after the incident was marked resolved. Several other US West users are reporting the same today.

Project: Foodly (ea81f764-8229-4e50-90cc-1f41d65781c1), environment production

Service: db-foodly (Postgres 18.6), volume postgres-volume, us-west2, 1000 MB (486 MB used)

Evidence:

Checkpoint logs: the longest single-file sync went from under 1.5 s (for weeks) to 13–29 s, with the same write load (~20 MB of WAL per hour).

Commits wait 10–16 s on IO:WalSync / LWLock:WALWrite, so every write in our app takes 5–30 s.

No locks or long transactions. Postgres has been up since Aug 21 (not restarted by the incident).

We have NOT restarted or redeployed the database, since other users report it didn't help and left their services stuck. Please escalate to an engineer: could you check the host behind this volume and move it to healthy storage, without deleting or recreating the volume? This is a restaurant POS in service right now. Thank you.

Awaiting User Response

1 Replies

Status changed to Awaiting Railway Response Railway • 2 days ago


chandrika
EMPLOYEE

2 days ago

Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE

Redeploying or restarting will not help here, we suggest holding off on both for now

Our infrastructure team is looking into it, and we'll update the incident as soon as we have more


Status changed to Awaiting User Response Railway • 2 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...