volume disk I/O ~1000x slower in US West — production checkout down
johncris7ads
HOBBYOP

2 days ago

URGENT — production checkout down. Please escalate to a human engineer.

This continues thread volume-disk-i-o-degradation-in-us-west-a-6503791d (your support AI escalated it to the team).

Project: nexopay-checkout, volume nexopay-checkout-volume (zd11280), US West. Since ~15:00 UTC Oct 3 (around incident 72DDHCC1), the volume is extremely slow. A restart and a redeploy (00:15 UTC, deployment efe4151c) did not help; it got worse.

Measured in a 60 s window (00:29:53–00:30:53 UTC): 62 reads with average latency 757.6 ms and 37 writes with average 1,600 ms. That is ~1.2 IOPS and 0.006 MB/s, about 0.04% of the 3,000 IOPS / 70 MB/s limit. The container spent 45.8 of 60 s fully stalled on I/O (io.pressure). We are not hitting any limit; per-operation latency is ~1,000x normal.

Impact: payment checkouts take over 60 s or fail; we are losing sales right now.

Please check the host/disk of this volume and move the volume to a healthy host as soon as possible. Do not delete any data.

Solved

2 Replies

Status changed to Awaiting Railway Response Railway • 2 days ago


chandrika
EMPLOYEE

2 days ago

Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE

Redeploying or restarting will not help here, we suggest holding off on both for now

Our infrastructure team is looking into it, and we'll update the incident as soon as we have more


Status changed to Awaiting User Response Railway • 2 days ago


chandrika

Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE Redeploying or restarting will not help here, we suggest holding off on both for now Our infrastructure team is looking into it, and we'll update the incident as soon as we have more

johncris7ads
HOBBYOP

2 days ago

Thanks for confirming. Update from our side: we resolved it by moving to a new volume.

Restarting and redeploying didn't help, as you said. The old volume stayed degraded: about 760 ms per read and 1.6 s per write, at only ~1 IOPS.

What fixed it: we stopped the service, copied the SQLite database to a freshly created volume in the same region (US West), attached the new volume to the service and started the same image. Downtime was about 6 minutes. On the new volume we see ~0.7 ms per read, ~18 ms per write and no I/O stall. The app is back to normal.

The old volume (nexopay-checkout-volume) is detached and kept untouched, in case your team wants to look into it. Please do not delete it. We will not restart or redeploy until the incident is resolved. Thanks!


Status changed to Awaiting Railway Response Railway • 2 days ago


Status changed to Solved Railway • 2 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...