2 days ago
Hi. Since about 00:00 UTC on Oct 3, 2026, writes to my service's volume have been very slow, and it is getting worse. It continued after the US West incident (72DDHCC1) was marked resolved at 19:33 UTC.
Project: merry-intuition (48b7cd42-2851-41aa-8910-def93bca6921)
Service: web (1089e7fe-eb1f-4429-bf12-f3e531065e27), production
Volume: web-volume (76fed889-41eb-4c29-af31-5a59f249c503), mounted at /data, 500 MB (~350 MB used), us-west2
What we see:
The app is one process using a SQLite database on the volume. A save writes a few KB to the database's log file (no forced disk sync).
Those small writes now take 1–4 seconds each, and opening the database file sometimes takes over 1 second. Until Oct 2 they took about 25–40 ms.
CPU is near idle (about 2% of one core). Read-only requests are still about 10 ms.
Our logs time each step: the waits are spent inside the operating system's write calls to the volume, not waiting on anything in our code.
Median response time for a save request: about 30 ms on Oct 1–2, then 80 ms, 170 ms, 580 ms and about 1,200 ms over Oct 3 (UTC).
Redeploying (17:17, 21:51 and 23:01 UTC) has not helped.
Could you check whether the host behind this volume is still degraded after the incident, and move the volume to a healthy host if so? Thank you.
Attachments
0 Replies
Status changed to Awaiting Railway Response Railway • 2 days ago
Status changed to Solved belac100 • 2 days ago