2 days ago
Privly needs reliable production saves without blocking user access. During the outage, requests intermittently returned Cloudflare 502 while the API waited for disk confirmation. Access is currently restored using SQLite WAL synchronous=NORMAL, a temporary durability compromise. Please investigate the affected persistent-volume host and recommend remediation that preserves customer data.
Resources:
Project: e8d2a923-59cf-4f00-a358-22077ea8041b
Service: privly / 8ccea99b-bb31-4895-b66a-79e02f1bc2be
Environment: production / a5cdfbab-75ed-4cb4-b138-1e446c8770f1
Volume: c9170c16-4a5f-4b1b-aa9c-1b720f96e35f, mount /data
Domains: https://www.privly.app and https://api.privly.app
Initial observed failures: 2026-10-03 23:11, 23:12 and 23:14 UTC, source commit 20834ded (before the subsequent Founder UI release).
Current successful deployment: 89c9072e-07ed-47dc-ae49-e2b4e00bb398, source ca454419 (2026-10-04 UTC).
Evidence:
Around 2026-10-04 01:06 UTC, an independent 4096-byte scratch-file probe measured:
/tmp write 0.037 ms; fsync 0.314 ms
/data write 0.011 ms; fsync 6669.054 ms
The probe used unique private temporary files, cleaned them up, and never opened customer records.
During the outage the API process blocked in kernel wait jbd2_log_wait_commit, syscall fsync, on /data/privly.sqlite-wal. Another observation showed folio_wait_bit_common. The volume had approximately 1.13 GB free. Recorded CPU and memory usage were far below configured limits.
Attempts and results:
We changed production WAL synchronous mode to NORMAL, preserved FULL as the application default, decoupled worker failures from API uptime and batched analytics bookkeeping. Public app/health requests then succeeded quickly, and recent observed logs had no requests over five seconds. The independent /data fsync probe still showed a 6.7-second delay after initial recovery. A further reduction in analytics commits is now deployed; user records and access remained unchanged.
Why support:
The AI assistant recommended contacting Railway because the independent persistent-volume probe remained slow after recovery. This isolates a disk-confirmation delay from application SQL, but does not establish its underlying platform cause. Our available tooling does not expose volume-host placement or a safe relocation operation. Other Central Station threads report similar persistent-volume latency on Oct 3; we have not confirmed a shared cause.
Please investigate host/volume latency, queueing, throttling and maintenance, and advise whether repair or relocation is appropriate. Keep customer data intact. Do not delete, restore or modify the database; coordinate any downtime. This request authorizes diagnosis and recommendations, not destructive changes.
Prepared with Railway's support template (v1).
1 Replies
Status changed to Awaiting Railway Response Railway • 2 days ago
2 days ago
Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE
Redeploying or restarting will not help here, we suggest holding off on both for now
Our infrastructure team is looking into it, and we'll update the incident as soon as we have more
Status changed to Awaiting User Response Railway • 2 days ago