Production outage: repeated fdatasync waits on /data volume after redeploy
alexborisov979-source
PROOP

2 days ago

Please investigate recurring storage flush stalls affecting SKVOZO production.

Project: SKVOZO

Project ID: 50768583-22aa-4838-8173-16816517c4f4

Service: skvozo-bot

Service ID: 1208eb57-c4d0-44d1-bd16-066bf8c752b2

Environment: production

Environment ID: d48257f9-622e-48b0-a7da-4602a7d29506

Region: sfo

Volume: skvozo-data

Volume ID: 54194eb0-9035-4375-a50d-31c02c81632e

Mount: /data

Previous deployment:

f626052a-ccf7-4367-a7df-111777ce8674

Replacement deployment:

5cc74f62-433a-402f-ae67-5c57ecd2dcb6

Commit:

40bc7e6ec1a5d66bb5a9c5c32f099ccf890ac7cf

Incident timeline (UTC):

  • Incident observed around 2026-10-03 23:56–2026-10-04 00:25.

  • Restart attempt around 00:41 failed with:

    “stacker-hooks precreate: exit status 2”

  • Redeploy of the same commit was initiated around 00:44.

  • Application started around 00:52 and the initial healthcheck passed.

  • HTTP 499 responses recurred by approximately 00:52:20–00:52:40.

Railway Agent reported repeated snapshots of Python PID 1 in D state, with the main thread waiting in:

submit_bio_wait

blkdev_issue_flush

ext4_sync_file

do_fsync

__x64_sys_fdatasync

Reported syscall:

75 0x8 0x0 0x0 0x0 0x0 0x0

This corresponds to fdatasync(8).

/proc/1/fdinfo/8:

pos: 0

flags: 02500002

mnt_id: 84923

ino: 15

/proc/1/mountinfo:

84923 84805 230:14928 / /data rw,relatime - ext4 /dev/zd14928 rw,discard,stripe=8192

The available Agent tools could not resolve FD 8 / inode 15 to a filename or capture the Python stack.

The application uses SQLite WAL with synchronous=FULL.

Synchronous SQLite operations currently share the main asyncio thread, so a prolonged storage flush blocks HTTP handling.

Impact:

  • Mini App becomes unavailable
  • GET / becomes unavailable
  • /health becomes unavailable
  • /ready becomes unavailable
  • /ready/fleet becomes unavailable
  • /bridge/health and /bridge/poll return 499/timeouts
  • Railway still reports the service as Online/SUCCESS

Please investigate:

  1. Resolve FD 8 / inode 15 to the actual file during a correlated blocked snapshot.
  2. Inspect host-side flush latency, queueing, and errors for this volume/device.
  3. Correlate storage telemetry with the outage and recurrence after redeploy.
  4. Determine whether the Railway volume/device is experiencing abnormal flush stalls.
  5. Recommend a data-preserving recovery and permanent remedy.

Please do not delete or replace the database/WAL, disable durability, or migrate/delete the volume without explicit approval.

Solved

0 Replies

Status changed to Awaiting Railway Response Railway • 2 days ago


Status changed to Solved alexborisov979-source • 2 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...