2 days ago
Please investigate recurring storage flush stalls affecting SKVOZO production.
Project: SKVOZO
Project ID: 50768583-22aa-4838-8173-16816517c4f4
Service: skvozo-bot
Service ID: 1208eb57-c4d0-44d1-bd16-066bf8c752b2
Environment: production
Environment ID: d48257f9-622e-48b0-a7da-4602a7d29506
Region: sfo
Volume: skvozo-data
Volume ID: 54194eb0-9035-4375-a50d-31c02c81632e
Mount: /data
Previous deployment:
f626052a-ccf7-4367-a7df-111777ce8674
Replacement deployment:
5cc74f62-433a-402f-ae67-5c57ecd2dcb6
Commit:
40bc7e6ec1a5d66bb5a9c5c32f099ccf890ac7cf
Incident timeline (UTC):
-
Incident observed around 2026-10-03 23:56–2026-10-04 00:25.
-
Restart attempt around 00:41 failed with:
“stacker-hooks precreate: exit status 2”
-
Redeploy of the same commit was initiated around 00:44.
-
Application started around 00:52 and the initial healthcheck passed.
-
HTTP 499 responses recurred by approximately 00:52:20–00:52:40.
Railway Agent reported repeated snapshots of Python PID 1 in D state, with the main thread waiting in:
submit_bio_wait
blkdev_issue_flush
ext4_sync_file
do_fsync
__x64_sys_fdatasync
Reported syscall:
75 0x8 0x0 0x0 0x0 0x0 0x0
This corresponds to fdatasync(8).
/proc/1/fdinfo/8:
pos: 0
flags: 02500002
mnt_id: 84923
ino: 15
/proc/1/mountinfo:
84923 84805 230:14928 / /data rw,relatime - ext4 /dev/zd14928 rw,discard,stripe=8192
The available Agent tools could not resolve FD 8 / inode 15 to a filename or capture the Python stack.
The application uses SQLite WAL with synchronous=FULL.
Synchronous SQLite operations currently share the main asyncio thread, so a prolonged storage flush blocks HTTP handling.
Impact:
- Mini App becomes unavailable
- GET / becomes unavailable
- /health becomes unavailable
- /ready becomes unavailable
- /ready/fleet becomes unavailable
- /bridge/health and /bridge/poll return 499/timeouts
- Railway still reports the service as Online/SUCCESS
Please investigate:
- Resolve FD 8 / inode 15 to the actual file during a correlated blocked snapshot.
- Inspect host-side flush latency, queueing, and errors for this volume/device.
- Correlate storage telemetry with the outage and recurrence after redeploy.
- Determine whether the Railway volume/device is experiencing abnormal flush stalls.
- Recommend a data-preserving recovery and permanent remedy.
Please do not delete or replace the database/WAL, disable durability, or migrate/delete the volume without explicit approval.
0 Replies
Status changed to Awaiting Railway Response Railway • 2 days ago
Status changed to Solved alexborisov979-source • 2 days ago