6 days ago
Our production Postgres stalls for 3–7 minutes each day during high user activity periods. The stalls start at the same times as the pgBackRest base backups that point-in-time recovery (PITR) runs. We would like to move these backups to an off-peak window, is it possible?
Limits: 10 vCPU, 14 GB memory, 50 GB volume. The database is about 21 GB.
PITR is on. A daily volume snapshot runs at 09:56 UTC; the snapshot does not cause stalls.
What happens
During each backup, simple queries hang for more than 15 s. Examples are single-row INSERTs, primary-key UPDATEs and small SELECTs. Our app timeout then cancels them, and the Postgres log shows ERROR: canceling statement due to user request.
Our web service and our worker service stall at the same minutes, so the cause is on the database side.
Most queries in the window still complete. Request volume is normal, so the load does not come from traffic.
Outside the backup windows we see no such cancellations.
Timeline (UTC), from the Postgres service logs
+--------+--------+-------------------+---------+--------------+
| Date | Backup | Start-end | Size | Our timeouts |
+--------+--------+-------------------+---------+--------------+
| Sep 22 | full | 12:33:55-12:41:21 | 19.9 GB | 12:35-12:39 |
| Sep 28 | diff | 10:55-10:59 | 17.1 GB | 10:55-10:58 |
| Sep 29 | diff | 11:00:27-11:05:02 | 17.1 GB | 11:01-11:03 |
| Sep 29 | full | 12:41:29-12:47:55 | 20.8 GB | 12:41-12:45 |
+--------+--------+-------------------+---------+--------------+
The days in between show the same pattern.
What we see in the metrics
CPU during the backups peaks at about 1.2 of 10 vCPU, so CPU is not the bottleneck.
Memory stays at the 14 GB limit. shared_buffers is 4 GB; the rest is the operating-system file cache.
The dashboard shows no disk read or write metric, so we cannot confirm I/O saturation.
Our assumption is that the backup reads 17–21 GB from the volume in 4–7 minutes, and that this saturates the volume's I/O or empties the file cache. Queries that need disk, and WAL flushes at commit, then wait.
Questions
Backup time. Can you fix the PITR base backup time at 03:00 UTC for this service? The start time now drifts about 5 minutes later each day and sometimes jumps (from about 12:52 on Sep 25 to 10:45 on Sep 26). It is now in our peak traffic. Our own offsite pg_dump runs at 03:32 UTC, and a backup at 03:00 would finish before it.
Disk limits. What are the IOPS and throughput limits of our volume? Is a faster volume tier available?
Throttling. Does Railway limit the I/O of the backup, for example with ionice or process-max? Can you limit it for our service?
Diff size. Each daily diff is 80–85% of a full backup. Does your image use pgBackRest block incremental backups (repo-block)? Could you turn it on for us?
6 Replies
Status changed to Awaiting Railway Response Railway • 6 days ago
6 days ago
The pgBackRest base backups that PITR takes run on a rolling schedule inside the Postgres container, a full backup weekly and a differential daily, and there is no setting that pins them to a clock time such as 03:00 UTC. The railway postgres pitr schedule commands only manage the daily, weekly and monthly volume backup schedules (your 09:56 UTC snapshot), not the PITR base backups.
Your Postgres volume is currently capped at 3,000 IOPS and 70 MB/s throughput, which is the Pro default. Your own figures of 17-21 GB read in 4-7 minutes work out to roughly 45-70 MB/s, so you can compare that directly against the cap.
On throttling and diff size, the PITR image does not expose options for I/O priority, pgBackRest process-max, or block incremental (repo-block) backups, so none of those can be changed from the service's configuration.
Status changed to Awaiting User Response brody • 6 days ago
6 days ago
Hi @brody, thanks for your response, it confirms the diagnosis, but I'm not sure what our options are to avoid making our system unresponsive to paying customers at the peak times.
- Our logs show the watcher starts each diff 24 h after the previous backup finished (Sep 27 end 10:54:58 → Sep 28 start 10:54:59; Sep 28 end 10:59:40 → Sep 29 start 11:00:27), and each full 7 days after the previous full finished. Is there a way we (or you) can trigger one diff and one full at 03:00 UTC to move the schedule?
For example, on Sep 26 the diff start jumped from about 12:52 to 10:45. What caused that? Would a container restart at 03:00 UTC also move the schedule?
Any other suggestions?
Status changed to Awaiting Railway Response Railway • 6 days ago
6 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 6 days ago
5 days ago
Hey — building on @brody’s note with what the official docs support (and what they don’t).
Schedule
Postgres PITR uses pgBackRest. Docs say it takes base backups on a rolling schedule: full weekly, differential daily (Point-in-Time Recovery). There is no published control to pin those to a clock time such as 03:00 UTC.
railway postgres pitr schedule manages volume backup schedules, not the PITR/pgBackRest base-backup clock (CLI postgres, Volume backups). That matches your 09:56 UTC volume snapshot being separate (and not stalling the same way).
Why stalls fit the published I/O caps
Volumes are fixed at 3,000 read IOPS and 3,000 write IOPS for all sizes and plans (Volumes reference → I/O). A multi‑GB base/diff backup shares that volume with OLTP, so short hangs during the window are consistent with I/O contention. Growing the volume does not raise those IOPS limits.
Documented options
- Keep volume snapshots on an off-peak schedule you control for snapshot restores (Volume backups).
- If you do not need continuous PITR, disable it (
railway postgres pitr disable/ Backups tab) and rely on volume backups — removes pgBackRest base backups, loses arbitrary-timestamp restore (PITR, CLI). - App-side: treat known backup windows as maintenance (retry/backoff, longer statement timeouts) until a schedule pin exists.
Undocumented (do not invent)
Docs do not publish a customer knob to move the next full+diff to 03:00 UTC, whether a container restart reliably shifts the rolling watcher, or pgBackRest I/O-priority tunables. For an off-peak one-shot or staff-side schedule move, ask Railway (@brody / Help Station) to confirm.
If continuous PITR is optional vs daily snapshots, that choice is the main self-serve lever docs give you today.
5 days ago
..
5 days ago
@brody, this is basically unusable for us. It's affecting users throughout the day for significant windows, and it looks like Railway doesn't have a solution to the problem and isn't willing to help. I guess the only option is to remove PITR and cross our fingers, which is not realistic for us long term. Having just sunk a load of time and effort into migrating all of our apps from Heroku over the past 6 months after hearing great things about Railway, this is extremely disappointing.
4 days ago
DOES ANYONE MONITOR ANY OF THE FOLLOW UPS HERE?
@brody we found what moves the schedule, in the Postgres service logs.
Every jump follows a failed WAL archive push:
ERROR: [082]: unable to push WAL file ... to the archive asynchronously after 1m second(s)
One minute later the watcher logs "gap-recovery: entering recovery" and takes an extra diff "to anchor restore point". The next
daily diff then starts 24 hours after that one finished.
- Sep 26 10:44 UTC: push failed, extra diff at about 10:45. This is the jump from 12:52 to 10:45 that we asked about.
- Sep 29 15:41 UTC: push failed, extra diff 15:43-15:46.
- Sep 30 07:42 UTC: push failed, extra diff 07:43-07:46. The daily diff has run at 07:46 (Oct 1) and 07:51 (Oct 2) since.
So the watcher can already take an out-of-schedule diff, and that diff re-anchors the daily clock. A redeploy does not: the
weekly full kept its rhythm across our Sep 19 deploy (Sep 15 12:26, Sep 22 12:33, Sep 29 12:41).
Two requests:
- Please trigger one diff at 03:00 UTC for this service, and the next weekly full at 03:00 UTC too. It is due about 12:48 UTC
on Oct 6, in the middle of an event day for us. If you cannot, please tell us whether it is safe for us to run "pgbackrest
backup --type=diff" in the container at 03:00 UTC ourselves.
- Why are the WAL pushes failing? There were three in five days, and each one caused an extra 13 GB backup during our working
day.



