2 days ago
URGENT — production checkout down. Please escalate to a human engineer.
This continues thread volume-disk-i-o-degradation-in-us-west-a-6503791d (your support AI escalated it to the team).
Project: nexopay-checkout, volume nexopay-checkout-volume (zd11280), US West. Since ~15:00 UTC Oct 3 (around incident 72DDHCC1), the volume is extremely slow. A restart and a redeploy (00:15 UTC, deployment efe4151c) did not help; it got worse.
Measured in a 60 s window (00:29:53–00:30:53 UTC): 62 reads with average latency 757.6 ms and 37 writes with average 1,600 ms. That is ~1.2 IOPS and 0.006 MB/s, about 0.04% of the 3,000 IOPS / 70 MB/s limit. The container spent 45.8 of 60 s fully stalled on I/O (io.pressure). We are not hitting any limit; per-operation latency is ~1,000x normal.
Impact: payment checkouts take over 60 s or fail; we are losing sales right now.
Please check the host/disk of this volume and move the volume to a healthy host as soon as possible. Do not delete any data.
2 Replies
Status changed to Awaiting Railway Response Railway • 2 days ago
2 days ago
Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE
Redeploying or restarting will not help here, we suggest holding off on both for now
Our infrastructure team is looking into it, and we'll update the incident as soon as we have more
Status changed to Awaiting User Response Railway • 2 days ago
chandrika
Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE Redeploying or restarting will not help here, we suggest holding off on both for now Our infrastructure team is looking into it, and we'll update the incident as soon as we have more
2 days ago
Thanks for confirming. Update from our side: we resolved it by moving to a new volume.
Restarting and redeploying didn't help, as you said. The old volume stayed degraded: about 760 ms per read and 1.6 s per write, at only ~1 IOPS.
What fixed it: we stopped the service, copied the SQLite database to a freshly created volume in the same region (US West), attached the new volume to the service and started the same image. Downtime was about 6 minutes. On the new volume we see ~0.7 ms per read, ~18 ms per write and no I/O stall. The app is back to normal.
The old volume (nexopay-checkout-volume) is detached and kept untouched, in case your team wants to look into it. Please do not delete it. We will not restart or redeploy until the incident is resolved. Thanks!
Status changed to Awaiting Railway Response Railway • 2 days ago
Status changed to Solved Railway • 2 days ago