7 hours ago
I'm seeing a sudden performance degradation on a Railway-hosted PostgreSQL 18.6 database.
The application began experiencing 10–30 second database-backed requests today. PostgreSQL checkpoint logs show a clear change beginning around 13:01 UTC:
- Before 13:01 UTC: checkpoint sync/flush times were typically 5–40 ms, with whole checkpoints under ~1.3 seconds.
- Starting around 13:01 UTC: sync/flush times increased to roughly 0.2–5.5 seconds, with some checkpoints taking up to ~24 seconds.
The application and database were running normally for hours before this change, with no deployment corresponding to the onset.
Railway metrics currently show very low PostgreSQL CPU usage, approximately 75–100 MB memory usage, and plenty of unused volume capacity. There have been no service restarts, deadlocks, database errors, or 5xx responses.
At idle, queries are currently completing in approximately 2–7 ms and there are no waiting locks or stuck transactions. During the degradation periods, however, normally fast requests queue for many seconds waiting for database connections/writes to complete.
This appears consistent with intermittent high write/fsync latency on the underlying PostgreSQL volume rather than CPU, memory, database locking, or application load.
Is there a way to check the underlying volume/storage I/O health for this PostgreSQL service, or determine whether the service is experiencing degraded storage performance on its host?
I found an existing Railway thread describing very similar intermittent WAL/fsync latency where migration to another region was suggested as a diagnostic step. Before moving a production database, I'd prefer to determine whether the underlying volume/host for this service is currently experiencing degraded I/O.
1 Replies
Status changed to Awaiting Railway Response Railway • about 7 hours ago
5 hours ago
We checked the storage behind your Postgres volume. The host it runs on has had sharply elevated disk write latency since about 13:00 UTC today, with average write wait times rising several-fold from their earlier baseline. That lines up with the checkpoint sync and flush times you measured.
Your volume itself is healthy and well under its capacity and I/O limits, and nothing in your database or application configuration is causing this. The slowdown comes from the shared disk on that host, not from your workload.
Status changed to Awaiting User Response Railway • about 5 hours ago