21 days ago
Hi Railway team,
Point-in-time recovery into a new service fails for us every time, and the restored service does not keep the source service's image. Details below.
Account / project
Plan: Hobby (upgraded 2026-09-14)
Project: ingenious-optimism (0239b1fd-b66d-42bb-afd2-9834334e54d2)
Environment: staging (2b54ac9c-2715-4b9d-9db1-85596679ea59)
Source service: pg-dict-probe — Postgres built from the official ghcr.io/railwayapp-templates/postgres-ssl:18 template with our own Dockerfile layered on top (deployed via railway up with POSTGRES_IMAGE=ghcr.io/railwayapp-templates/postgres-ssl:18 and RAILWAY_DOCKERFILE_PATH=Dockerfile). PITR status: enabled, bucket wired, WAL archiver healthy. Volume 209 MB / 500 MB.
CLI: railway 5.41.2
Steps to reproduce (done twice, same result both times)
railway postgres pitr restore -s pg-dict-probe -e staging --at 10m --new-service-name pg-dict-restore-test --yes Workflow: createServiceFromPITR/0239b1fd-b66d-42bb-afd2-9834334e54d2/772334ef-b663-49ca-9e11-4068e4c1c88a/s5YuE9V87CbjZrsJ1lfRQ
railway postgres pitr restore -s pg-dict-probe -e staging --at 2026-09-13T22:30:00Z --new-service-name pg-dict-restore-2 --yes Workflow: createServiceFromPITR/0239b1fd-b66d-42bb-afd2-9834334e54d2/772334ef-b663-49ca-9e11-4068e4c1c88a/WEGz_Osio2m2fQcckA0mo
What happens
Recovery reaches a consistent state and the stop point, then the container crashes:
LOG: consistent recovery state reached at 10/B8000098
LOG: recovery stopping after commit of transaction 780, time 2026-09-13 22:12:05.665993+00
ERROR: [042]: unable to get 0000000100000011000000D1: ... [FileWriteError] unable to write '/var/lib/postgresql/data/pgdata/pg_wal/RECOVERYXLOG
FATAL: could not write to file "pg_wal/xlogtemp.73": No space left on device
LOG: startup process (PID 73) exited with exit code 1
Deployment status: CRASHED. The restored volume shows 4 MB / 5000 MB used, and the source database is only 209 MB, so we don't think the volume itself is full. Our guess is that the space runs out in some temporary storage during restore, or the volume size is applied after provisioning.
The restored service's source is the plain official image ghcr.io/railwayapp-templates/postgres-ssl:18, not the image the source service actually runs. Our image adds a text-search dictionary and a Postgres extension that our schema depends on. A database restored onto the plain image would come up without them, and our migrations and search would fail.
(Both test services and their volumes were deleted after the test.)
Questions
Why does PITR restore fail with No space left on device when the new volume is nearly empty? Is there a size limit on temporary storage during restore on Hobby, and how do we get around it?
Can a PITR restore create the new service on the same build/image as the source service (a service deployed from a Dockerfile via railway up), rather than the official template image?
If not: what is the supported way to switch a restored service to our own image (redeploy with railway up onto the restored service) without losing the recovered data on the volume?
Our production Postgres (production environment, same setup, PITR enabled) would hit the same path in an incident. Is there anything we should configure now so a production restore works?
Thanks!
1 Replies
Status changed to Awaiting Railway Response Railway • 21 days ago
21 days ago
The crash is the restored volume filling during WAL replay. A restore lays down the base backup then replays every archived WAL segment up to your target, and the volume must hold both at once. Peak usage is driven by how much WAL sits between the nearest base backup and the target, not by the database size. The two levers are restoring to a target closer to a recent base backup (less WAL to replay) or using a plan with a higher volume size ceiling. The source service is never touched by a restore.
On the image: the restore creates the new service with the same image as the source, meaning whatever the source service's configured source is. You can check your source service's configured source in its settings. Once recovery finishes successfully, the restored service is an ordinary service, so you can deploy your Dockerfile onto it with railway up and the volume keeps the recovered data through the redeploy. Use an image built on the same Postgres major the restore landed on.
The same applies to production - no additional configuration is needed beyond ensuring enough volume headroom for WAL replay.
Status changed to Awaiting User Response Railway • 21 days ago
14 days ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • 14 days ago