5 days ago
Hi Railway team,
I’m testing PITR recovery on a production PostgreSQL 18.6 service using the Railway Postgres SSL template.
We reproduced the same restore failure twice on isolated PITR sibling services.
The important finding is that this does not appear to be caused by a long WAL replay.
In the latest controlled test:
-
PITR base backup stop time:
2026-09-29 15:04:18Z -
Restore target:
2026-09-29 15:04:19Z -
PostgreSQL reached the requested target almost immediately:
recovery stopping before commit of transaction 771, time 2026-09-29 15:04:19.403329+00 -
redo doneoccurred after roughly two WAL segments -
However,
archive-getcontinued fetching about 285 contiguous 16 MiB WAL segments -
pg_walgrew from ~545 MiB to ~4.5 GiB -
The restore volume filesystem (
/var/lib/postgresql/data, ~4.6 GiB usable) reached 100% -
Inode usage remained ~1%
-
The container root filesystem had plenty of free space
-
/var/spool/pgbackrestwas only ~4 KiB
PostgreSQL then failed with:
FileWriteError ... pg_wal/RECOVERYXLOG: [28] No space left on device
and:
FATAL: could not write to file "pg_wal/xlogtemp...": No space left on device
Railway still showed the restore deployment as SUCCESS, while PostgreSQL was not accepting connections and remained in archive recovery.
The same approximate failure point — ~285 WAL segments / ~4.5 GiB — also occurred in an earlier restore targeting a timestamp approximately 18 hours after the eligible base backup.
So the second test appears to rule out replay length as the main driver: a restore that required only about two WAL segments to reach its target still fetched about 285 segments and filled the restore volume before promotion.
Could you help clarify the following?
- Why does a PITR restore fetch ~285 WAL segments when the recovery target is reached after roughly two?
- Are WAL segments fetched beyond the recovery target expected to remain in
pg_waluntil promotion? - Can
recovery_prefetchor another Railway restore mechanism cause this read-ahead? - Can
recovery_prefetchbe disabled or controlled during a Railway PITR restore? - How is the sibling restore volume sized?
- Is the restore volume expected to accommodate all prefetched WAL?
- Is this a known issue with mostly idle PostgreSQL databases using
archive_timeout=1min? - Why does the restore wrapper use
max_wal_size=512MBandcheckpoint_timeout=30sif no restartpoint/checkpoint occurs before the filesystem fills? - What exactly does
maxRestoreTimerepresent for PostgreSQL PITR? - Is Railway deployment status
SUCCESSexpected while the restored PostgreSQL process is crash-looping or not ready? - What mitigation would you recommend: restore-volume sizing,
recovery_prefetch,archive_timeout, a template/image change, or a platform fix?
The source database remained untouched and healthy during both drills. Both restore siblings were isolated and later deleted.
I can provide sanitized restore logs, df / pg_wal measurements, PITR status output, and deployment IDs privately if useful.
Thanks.
3 Replies
5 days ago
We checked the restore volume from your latest drill. It was created at 5 GB, which is the larger of the source volume's size and the default size for a new volume, and it was essentially full when the restore stopped. 5 GB is also the largest volume the Hobby plan allows. A PITR restore has to hold the base data plus the write-ahead log it pulls from the archive at the same time, and here that outgrew the volume. Your source database was not affected.
Restoring to a target closer to a backup reduces how much room a restore needs, but your 15:04:19 test was already one second after the backup and still filled the volume, so that option is used up. The other option is a bigger restore volume. A restored service continues its recovery when you redeploy it after its volume is grown. Your restore volume is already at the Hobby maximum, so on this plan it cannot grow any further. A plan with a higher volume size limit is what would give the restore more room.
Both leftover restore volumes are already scheduled for deletion, so you don't need to clean anything up.
Status changed to Awaiting User Response Railway • 5 days ago
Railway
We checked the restore volume from your latest drill. It was created at 5 GB, which is the larger of the source volume's size and the default size for a new volume, and it was essentially full when the restore stopped. 5 GB is also the largest volume the Hobby plan allows. A PITR restore has to hold the base data plus the write-ahead log it pulls from the archive at the same time, and here that outgrew the volume. Your source database was not affected. Restoring to a target closer to a backup reduces how much room a restore needs, but your 15:04:19 test was already one second after the backup and still filled the volume, so that option is used up. The other option is a bigger restore volume. A restored service continues its recovery when you redeploy it after its volume is grown. Your restore volume is already at the Hobby maximum, so on this plan it cannot grow any further. A plan with a higher volume size limit is what would give the restore more room. Both leftover restore volumes are already scheduled for deletion, so you don't need to clean anything up.
5 days ago
Thanks, this clarifies the failure.
Before we run another recovery drill, could you confirm the recommended procedure?
Our source volume is currently 5 GB. If we temporarily move to a plan with a higher volume limit, would you recommend:
starting a new PITR restore with the source still at 5 GB,
increasing only the restored sibling volume after it is created,
redeploying that restored service so recovery resumes,
rather than increasing the production source volume?
Also, based on the behaviour you observed, what restore-volume size would you recommend for this database — for example 20 GB, 30 GB, or another value — so that a restore can safely cover the normal interval between PITR backups?
We want to avoid changing the source volume solely for the restore drill if growing the isolated sibling is the intended recovery procedure.
Thanks.
Status changed to Awaiting Railway Response Railway • 5 days ago
5 days ago
Yes, that is the intended procedure. The restored service gets its own volume, and growing that volume and then redeploying the restored service continues the recovery from where it stopped, so you don't need to change the production source volume for the drill. The source stays untouched throughout.
We can't give you a specific size to pick. How much room a restore needs depends on how much write-ahead log sits between the backup and your target, not on how big the database is. Before you start a restore, the Backups tab in the dashboard estimates how much room it will need and warns you if the volume is too small. Use that estimate to size the restored volume.
Status changed to Awaiting User Response Railway • 5 days ago