2 days ago
Hi, since the US West incident on Oct 3 (UTC), our service has frequent
write stalls on its volume.
- Project: empowering-healing
- Service: worker (region: sfo / US West, with a 5 GB volume)
- App: a Discord bot using SQLite on the volume
What we see:
-
Reads/writes on the volume block for about 5-15 seconds, roughly every
10-20 minutes.
-
Examples (2026-10-04, UTC): 03:01, 03:19, 03:40, 03:44, 03:50, 03:58,
04:05, 04:23, 04:29, 04:38, 04:44, 05:02, 05:07, 05:29, 05:33, 05:44, 06:00
-
CPU stays near 0% and memory around 0.3 GB, so it doesn't look like a
resource limit on our side.
-
Before the incident this happened only occasionally; it is much more
frequent now.
-
We redeployed at 02:18 UTC on Oct 4, but it continues.
Could you check whether the host or volume we are on is degraded, or move
us to a healthy host?
Thank you.
3 Replies
Status changed to Awaiting Railway Response Railway • 1 day ago
a day ago
Update: I moved the service to US East (us-east4) on Oct 4, 23:39 UTC.
The volume migration took about a minute.
Before the move (sfo), even tiny commits took 1-4.6 seconds with nothing
else queued, which points to the volume itself. So far after the move,
there are no slow commits. I'll keep monitoring and report back.
It would still be good to know whether the sfo host/volume was degraded,
in case others are affected.
10 hours ago
It has improved after moving the service to US East (us-east4) on Oct 4,
23:39 UTC.
In sfo, small commits took 1-15 seconds every 10-20 minutes during the day.
In US East, over the first ~10 hours there was only one cluster of slow
commits: at Oct 5, 09:52-09:53 UTC, commits took 2.7s, 2.6s and 11.7s with
nothing else queued. Otherwise commits are fast.
So the move fixed most of it, but there are still rare multi-second stalls
on the US East volume as well. Is this kind of occasional fsync latency
expected on volumes, or could it also be checked on your side?
We're also planning to reduce fsyncs on our side (SQLite WAL mode).
an hour ago
On sfo: we had an incident affecting attached storage performance for some users in US West on Oct 4, between roughly 02:47 and 06:23 UTC, and every stall time you listed falls inside that window. Full timeline: Slow attached storage for some users in US West
On US East: the infrastructure your volume now runs on had a short burst of elevated write latency in the 09:50-09:55 UTC window on Oct 5, which lines up with the slow commits you saw at 09:52-09:53 UTC. Write latency was at normal low levels before and after that burst, and nothing in your service's own resource usage explains it.
Your volume is currently placed on healthy infrastructure in US East with the standard IOPS and bandwidth limits.
Status changed to Awaiting User Response Railway • about 1 hour ago