PostgreSQL PITR restore fills pg_wal to 100% and fails promotion even for a 2-segment replay
francescopaolodamico-dev
PROOP

5 days ago

Hi Railway team,

I’m testing PITR recovery on a production PostgreSQL 18.6 service using the Railway Postgres SSL template.

We reproduced the same restore failure twice on isolated PITR sibling services.

The important finding is that this does not appear to be caused by a long WAL replay.

In the latest controlled test:

  • PITR base backup stop time: 2026-09-29 15:04:18Z

  • Restore target: 2026-09-29 15:04:19Z

  • PostgreSQL reached the requested target almost immediately:

    recovery stopping before commit of transaction 771, time 2026-09-29 15:04:19.403329+00

  • redo done occurred after roughly two WAL segments

  • However, archive-get continued fetching about 285 contiguous 16 MiB WAL segments

  • pg_wal grew from ~545 MiB to ~4.5 GiB

  • The restore volume filesystem (/var/lib/postgresql/data, ~4.6 GiB usable) reached 100%

  • Inode usage remained ~1%

  • The container root filesystem had plenty of free space

  • /var/spool/pgbackrest was only ~4 KiB

PostgreSQL then failed with:

FileWriteError ... pg_wal/RECOVERYXLOG: [28] No space left on device

and:

FATAL: could not write to file "pg_wal/xlogtemp...": No space left on device

Railway still showed the restore deployment as SUCCESS, while PostgreSQL was not accepting connections and remained in archive recovery.

The same approximate failure point — ~285 WAL segments / ~4.5 GiB — also occurred in an earlier restore targeting a timestamp approximately 18 hours after the eligible base backup.

So the second test appears to rule out replay length as the main driver: a restore that required only about two WAL segments to reach its target still fetched about 285 segments and filled the restore volume before promotion.

Could you help clarify the following?

  1. Why does a PITR restore fetch ~285 WAL segments when the recovery target is reached after roughly two?
  2. Are WAL segments fetched beyond the recovery target expected to remain in pg_wal until promotion?
  3. Can recovery_prefetch or another Railway restore mechanism cause this read-ahead?
  4. Can recovery_prefetch be disabled or controlled during a Railway PITR restore?
  5. How is the sibling restore volume sized?
  6. Is the restore volume expected to accommodate all prefetched WAL?
  7. Is this a known issue with mostly idle PostgreSQL databases using archive_timeout=1min?
  8. Why does the restore wrapper use max_wal_size=512MB and checkpoint_timeout=30s if no restartpoint/checkpoint occurs before the filesystem fills?
  9. What exactly does maxRestoreTime represent for PostgreSQL PITR?
  10. Is Railway deployment status SUCCESS expected while the restored PostgreSQL process is crash-looping or not ready?
  11. What mitigation would you recommend: restore-volume sizing, recovery_prefetch, archive_timeout, a template/image change, or a platform fix?

The source database remained untouched and healthy during both drills. Both restore siblings were isolated and later deleted.

I can provide sanitized restore logs, df / pg_wal measurements, PITR status output, and deployment IDs privately if useful.

Thanks.

Awaiting User Response

3 Replies

Railway
BOT

5 days ago

We checked the restore volume from your latest drill. It was created at 5 GB, which is the larger of the source volume's size and the default size for a new volume, and it was essentially full when the restore stopped. 5 GB is also the largest volume the Hobby plan allows. A PITR restore has to hold the base data plus the write-ahead log it pulls from the archive at the same time, and here that outgrew the volume. Your source database was not affected.

Restoring to a target closer to a backup reduces how much room a restore needs, but your 15:04:19 test was already one second after the backup and still filled the volume, so that option is used up. The other option is a bigger restore volume. A restored service continues its recovery when you redeploy it after its volume is grown. Your restore volume is already at the Hobby maximum, so on this plan it cannot grow any further. A plan with a higher volume size limit is what would give the restore more room.

Both leftover restore volumes are already scheduled for deletion, so you don't need to clean anything up.


Status changed to Awaiting User Response Railway • 5 days ago


Railway

We checked the restore volume from your latest drill. It was created at 5 GB, which is the larger of the source volume's size and the default size for a new volume, and it was essentially full when the restore stopped. 5 GB is also the largest volume the Hobby plan allows. A PITR restore has to hold the base data plus the write-ahead log it pulls from the archive at the same time, and here that outgrew the volume. Your source database was not affected. Restoring to a target closer to a backup reduces how much room a restore needs, but your 15:04:19 test was already one second after the backup and still filled the volume, so that option is used up. The other option is a bigger restore volume. A restored service continues its recovery when you redeploy it after its volume is grown. Your restore volume is already at the Hobby maximum, so on this plan it cannot grow any further. A plan with a higher volume size limit is what would give the restore more room. Both leftover restore volumes are already scheduled for deletion, so you don't need to clean anything up.

francescopaolodamico-dev
PROOP

5 days ago

Thanks, this clarifies the failure.

Before we run another recovery drill, could you confirm the recommended procedure?

Our source volume is currently 5 GB. If we temporarily move to a plan with a higher volume limit, would you recommend:

starting a new PITR restore with the source still at 5 GB,

increasing only the restored sibling volume after it is created,

redeploying that restored service so recovery resumes,

rather than increasing the production source volume?

Also, based on the behaviour you observed, what restore-volume size would you recommend for this database — for example 20 GB, 30 GB, or another value — so that a restore can safely cover the normal interval between PITR backups?

We want to avoid changing the source volume solely for the restore drill if growing the isolated sibling is the intended recovery procedure.

Thanks.


Status changed to Awaiting Railway Response Railway • 5 days ago


Railway
BOT

5 days ago

Yes, that is the intended procedure. The restored service gets its own volume, and growing that volume and then redeploying the restored service continues the recovery from where it stopped, so you don't need to change the production source volume for the drill. The source stays untouched throughout.

We can't give you a specific size to pick. How much room a restore needs depends on how much write-ahead log sits between the backup and your target, not on how big the database is. Before you start a restore, the Backups tab in the dashboard estimates how much room it will need and warns you if the volume is too small. Use that estimate to size the restored volume.


Status changed to Awaiting User Response Railway • 5 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...