2 days ago
Our production Postgres has had severe disk sync latency since ~12:00 UTC on Oct 3, starting with incident 72DDHCC1, and it got worse after the incident was marked resolved. Several other US West users are reporting the same today.
Project: Foodly (ea81f764-8229-4e50-90cc-1f41d65781c1), environment production
Service: db-foodly (Postgres 18.6), volume postgres-volume, us-west2, 1000 MB (486 MB used)
Evidence:
Checkpoint logs: the longest single-file sync went from under 1.5 s (for weeks) to 13–29 s, with the same write load (~20 MB of WAL per hour).
Commits wait 10–16 s on IO:WalSync / LWLock:WALWrite, so every write in our app takes 5–30 s.
No locks or long transactions. Postgres has been up since Aug 21 (not restarted by the incident).
We have NOT restarted or redeployed the database, since other users report it didn't help and left their services stuck. Please escalate to an engineer: could you check the host behind this volume and move it to healthy storage, without deleting or recreating the volume? This is a restaurant POS in service right now. Thank you.
1 Replies
Status changed to Awaiting Railway Response Railway • 2 days ago
2 days ago
Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE
Redeploying or restarting will not help here, we suggest holding off on both for now
Our infrastructure team is looking into it, and we'll update the incident as soon as we have more
Status changed to Awaiting User Response Railway • 2 days ago
