Frequent volume write stalls (5-15s) on SQLite since the US West incident
kaijpn
HOBBYOP

2 days ago

Hi, since the US West incident on Oct 3 (UTC), our service has frequent

write stalls on its volume.

  • Project: empowering-healing
  • Service: worker (region: sfo / US West, with a 5 GB volume)
  • App: a Discord bot using SQLite on the volume

What we see:

  • Reads/writes on the volume block for about 5-15 seconds, roughly every

    10-20 minutes.

  • Examples (2026-10-04, UTC): 03:01, 03:19, 03:40, 03:44, 03:50, 03:58,

    04:05, 04:23, 04:29, 04:38, 04:44, 05:02, 05:07, 05:29, 05:33, 05:44, 06:00

  • CPU stays near 0% and memory around 0.3 GB, so it doesn't look like a

    resource limit on our side.

  • Before the incident this happened only occasionally; it is much more

    frequent now.

  • We redeployed at 02:18 UTC on Oct 4, but it continues.

Could you check whether the host or volume we are on is degraded, or move

us to a healthy host?

Thank you.

Awaiting User Response

3 Replies

Status changed to Awaiting Railway Response Railway • 1 day ago


kaijpn
HOBBYOP

a day ago

Update: I moved the service to US East (us-east4) on Oct 4, 23:39 UTC.

The volume migration took about a minute.

Before the move (sfo), even tiny commits took 1-4.6 seconds with nothing

else queued, which points to the volume itself. So far after the move,

there are no slow commits. I'll keep monitoring and report back.

It would still be good to know whether the sfo host/volume was degraded,

in case others are affected.


kaijpn
HOBBYOP

10 hours ago

It has improved after moving the service to US East (us-east4) on Oct 4,

23:39 UTC.

In sfo, small commits took 1-15 seconds every 10-20 minutes during the day.

In US East, over the first ~10 hours there was only one cluster of slow

commits: at Oct 5, 09:52-09:53 UTC, commits took 2.7s, 2.6s and 11.7s with

nothing else queued. Otherwise commits are fast.

So the move fixed most of it, but there are still rare multi-second stalls

on the US East volume as well. Is this kind of occasional fsync latency

expected on volumes, or could it also be checked on your side?

We're also planning to reduce fsyncs on our side (SQLite WAL mode).


an hour ago

On sfo: we had an incident affecting attached storage performance for some users in US West on Oct 4, between roughly 02:47 and 06:23 UTC, and every stall time you listed falls inside that window. Full timeline: Slow attached storage for some users in US West

On US East: the infrastructure your volume now runs on had a short burst of elevated write latency in the 09:50-09:55 UTC window on Oct 5, which lines up with the slow commits you saw at 09:52-09:53 UTC. Write latency was at normal low levels before and after that burst, and nothing in your service's own resource usage explains it.

Your volume is currently placed on healthy infrastructure in US East with the standard IOPS and bandwidth limits.


Status changed to Awaiting User Response Railway • about 1 hour ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...