a month ago
Our Postgres service is severely degraded since ~23:45 UTC on 2026-08-28, matching
the timeline of your active "Deployments slow to start / individual host" incident
and the other Postgres reports on Station tonight.
Project: insightful-courage
Environment: production
Service: Postgres (us-west2)
Evidence gathered from inside the containers:
-
Host loadavg is 400+ (read via /proc/loadavg from inside our container).
-
Checkpoints of <1.5 MB of data are taking 196-550 seconds:
checkpoint complete: wrote 83 buffers ... write=69.9s, sync=18.8s, total=196.2s
checkpoint complete: wrote 270 buffers ... write=236.9s, sync=43.5s, total=550.8s
(single-file fsyncs up to 24.8 s)
-
A cold-cache single-row SELECT takes ~8 s, while SELECT 1 on a warm connection
is 58 ms - so it is storage I/O, not network or app-side.
-
Volume is only 19% full (862 MB / 4.6 GB). No deploys on our side in the last
24 h, connection count is normal (~25 of 500).
-
App impact: requests queue behind WAL writes (user-facing latencies of 70-80 s,
lock convoys on hot rows). Yesterday 2026-08-23 04:22 UTC this same service also
had an unclean restart ("database system was interrupted"), possibly same host.
Could you check whether our Postgres volume sits on the affected host and migrate
it to a healthy one? Happy to run any diagnostics you need.
6 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
Seeing the same issue but in MySQL.
Did you find any way to fix it? I'm trying to escalate with the Railway team. Current declared incident doesn't seem to acknowledge this problem.
a month ago
Same issue here
ricardomacario
Seeing the same issue but in MySQL. Did you find any way to fix it? I'm trying to escalate with the Railway team. Current declared incident doesn't seem to acknowledge this problem.
a month ago
Thanks for chiming in, that's a useful data point. You seeing the same symptoms on MySQL confirms this is host-level storage degradation, not anything database-specific. Adding to the record, our region is us-west2 is yours too? I think that would help Railway narrow down the sick hosts. Still ongoing on our side as of now, latest Postgres checkpoint total=119.131 s for aprox 1 MB of data (healthy baseline for this service is <30 s), and user-facing requests are still hitting 25-30 s. It briefly improved around 00:39 UTC (down to about 105 s) and then regressed to 300 s, so whatever mitigation shipped for the deployment-queue incident didn't reach the storage side.
a month ago
Hey! Could you share if you're connecting to this database over the public or private network, and if you're seeing any recovery as of now?
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
Private networking, the app connects via postgres.railway.internal:5432. Note that our main evidence is independent of the network path though, the checkpoint timings come from the Postgres server's own logs (write/sync phases, i.e. local disk fsync), so this looks like storage I/O on the host rather than networking.
Partial recovery in progress, not complete. Checkpoint totals over the last 45 min: 300s (00:53) -> 142s -> 119s -> 69s (01:14) -> 88s (01:19). Healthy baseline for this service is under 30 seconds. Worst user-facing requests improved from 70-80 sec at peak to ~15-18 sec now. Load average inside the container was 400+ at peak; happy to re-measure if useful.
So, it's better since ~01:05 UTC but still aprox 2-3x degraded, with some oscillation. Can you confirm whether the host our Postgres volume lives on is the one from incident 8GL2R2U5, and whether a volume migration is on the table if it doesn't fully recover?
Status changed to Awaiting Railway Response Railway • about 1 month ago
a month ago
We've isolated the cause to our storage layer -- https://status.railway.com/incident/Z5Y3WO06. Apologies for the disruption!
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago