Production /data fsync takes 6.7 seconds and stalls Privly HTTP requests
giresskenne
HOBBYOP

2 days ago

Privly needs reliable production saves without blocking user access. During the outage, requests intermittently returned Cloudflare 502 while the API waited for disk confirmation. Access is currently restored using SQLite WAL synchronous=NORMAL, a temporary durability compromise. Please investigate the affected persistent-volume host and recommend remediation that preserves customer data.

Resources:

Project: e8d2a923-59cf-4f00-a358-22077ea8041b

Service: privly / 8ccea99b-bb31-4895-b66a-79e02f1bc2be

Environment: production / a5cdfbab-75ed-4cb4-b138-1e446c8770f1

Volume: c9170c16-4a5f-4b1b-aa9c-1b720f96e35f, mount /data

Domains: https://www.privly.app and https://api.privly.app

Initial observed failures: 2026-10-03 23:11, 23:12 and 23:14 UTC, source commit 20834ded (before the subsequent Founder UI release).

Current successful deployment: 89c9072e-07ed-47dc-ae49-e2b4e00bb398, source ca454419 (2026-10-04 UTC).

Evidence:

Around 2026-10-04 01:06 UTC, an independent 4096-byte scratch-file probe measured:

/tmp write 0.037 ms; fsync 0.314 ms

/data write 0.011 ms; fsync 6669.054 ms

The probe used unique private temporary files, cleaned them up, and never opened customer records.

During the outage the API process blocked in kernel wait jbd2_log_wait_commit, syscall fsync, on /data/privly.sqlite-wal. Another observation showed folio_wait_bit_common. The volume had approximately 1.13 GB free. Recorded CPU and memory usage were far below configured limits.

Attempts and results:

We changed production WAL synchronous mode to NORMAL, preserved FULL as the application default, decoupled worker failures from API uptime and batched analytics bookkeeping. Public app/health requests then succeeded quickly, and recent observed logs had no requests over five seconds. The independent /data fsync probe still showed a 6.7-second delay after initial recovery. A further reduction in analytics commits is now deployed; user records and access remained unchanged.

Why support:

The AI assistant recommended contacting Railway because the independent persistent-volume probe remained slow after recovery. This isolates a disk-confirmation delay from application SQL, but does not establish its underlying platform cause. Our available tooling does not expose volume-host placement or a safe relocation operation. Other Central Station threads report similar persistent-volume latency on Oct 3; we have not confirmed a shared cause.

Please investigate host/volume latency, queueing, throttling and maintenance, and advise whether repair or relocation is appropriate. Keep customer data intact. Do not delete, restore or modify the database; coordinate any downtime. This request authorizes diagnosis and recommendations, not destructive changes.

Prepared with Railway's support template (v1).

Awaiting User Response

1 Replies

Status changed to Awaiting Railway Response Railway • 2 days ago


chandrika
EMPLOYEE

2 days ago

Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE

Redeploying or restarting will not help here, we suggest holding off on both for now

Our infrastructure team is looking into it, and we'll update the incident as soon as we have more


Status changed to Awaiting User Response Railway • 2 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...