Postgres WAL sync stalls: 29-second disk syncs causing OAuth and pack failures
dolphstar
HOBBYOP

2 days ago

Our production WikiRivals application is experiencing intermittent PostgreSQL storage stalls. Please inspect the host and storage latency for the selected Postgres service during 2026-10-04 00:35–01:05 UTC.

Observed evidence:

  • Live pg_stat_activity showed API queries waiting on IO/WalSync and LWLock/WALWrite, including an auth query stalled for more than 9 seconds. No row-lock blockers were present.
  • Checkpoint at 01:04:24 UTC: write=43.443 s, sync=29.452 s, total=104.746 s; 299 buffers, 71 sync files, longest individual sync=5.777 s. Earlier sync phases reached 17.632 s.
  • OAuth callback timed out with HTTP 503 after 10 seconds. Pack-opening transactions hit statement timeouts and failed after 9–29 seconds. A bounded read-only diagnostic query also timed out.
  • Normal catalog checks finish in approximately 13–16 ms between stalls. The same application pack flow completes locally in 58–136 ms against a cloned catalog.
  • CPU was low: Postgres peak about 0.065 vCPU; memory peak about 0.767 GB of 8 GB; volume about 2.47 GB used of 5 GB.
  • PITR is enabled and healthy. fsync and synchronous_commit remain enabled.
  • A saved-opening retry and a later new opening succeeded, but slow sign-ins and storage stalls continued afterward.

Please confirm whether the underlying storage host is degraded and whether the volume needs migration to healthy storage. The affected Postgres service is selected in this support form; infrastructure identifiers are omitted from the public message.

No restart, redeploy, durability-setting change, or data migration has been performed in response to this incident.

Solved

1 Replies

Status changed to Awaiting Railway Response Railway • 2 days ago


chandrika
EMPLOYEE

2 days ago

Confirming that we're seeing this on our side, have called an incident here and are looking into this: https://status.railway.com/incident/A0O39CGE

Redeploying or restarting will not help here, we suggest holding off on both for now

Our infrastructure team is looking into it, and we'll update the incident as soon as we have more


Status changed to Awaiting User Response Railway • 2 days ago


Status changed to Solved dolphstar • 2 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...