11 days ago
WAL files archived to S3 successfully but not being deleted from local volume, causing disk to fill.
Service: Postgres
Environment: production
Volume: postgres-volume-1 (50 GB)
Current Situation:
- Volume disk usage: 1.17 GB, growing ~0.02 GB/hour (~0.5 GB/day)
- At current rate, volume fills in ~83 days
- Database is empty (no user data), only WAL generation from normal operation
What's Working:
- WAL archiving to S3 is successful — logs show "archive-push command end: completed successfully" for every file
- pgBackRest async archiving is configured and running
- S3 bucket: postgres-bucket-1
- No archiving failures (pgBackRest watcher shows failed=0)
What's Broken:
- Local WAL files are NOT being cleaned up after archiving
- Checkpoint logs consistently show: "0 WAL file(s) added, 0 removed, 5 recycled"
- New WAL segments keep being created but old archived ones are never deleted from /var/lib/postgresql/data/pg_wal
- Files are successfully pushed to S3 but persist locally indefinitely
Why This Matters:
- Need to keep WAL backups enabled (can't disable archiving)
- Volume will fill in ~83 days if growth continues
- This appears to be a pgBackRest template/configuration issue where async archiving succeeds but local cleanup is broken
Can this be fixed in the pgBackRest configuration, or is there a known issue with the current Railway Postgres template?
2 Replies
11 days ago
Having looked into this, the issue appears to be in your application code or configuration rather than the Railway platform itself, which puts it outside what Railway support can resolve directly.
This is exactly the kind of problem the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.
Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.
- Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
- Keep it private and close the thread - Nothing becomes public and the thread closes.
Status changed to Awaiting User Response Railway • 11 days ago
11 days ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • 11 days ago
11 days ago
This is usually PostgreSQL's WAL retention/recycling policy, not pgBackRest failing to delete a local file. pgBackRest's archive-push copies completed segments to the archive; PostgreSQL decides when segments in the WAL directory are no longer needed and then removes or recycles them at checkpoints. Your "0 removed, 5 recycled" line is consistent with normal recycling and does not by itself show a leak. An empty database still writes WAL.
To distinguish a bounded pool of recycled segments from real unbounded growth, please collect the following twice, several hours apart:
SHOW server_version;
SHOW min_wal_size;
SHOW max_wal_size;
SHOW wal_keep_size;
SHOW max_slot_wal_keep_size;
SELECT * FROM pg_stat_archiver;
SELECT slot_name, slot_type, active, restart_lsn FROM pg_replication_slots;
SELECT count(*) AS segments, pg_size_pretty(sum(size)::bigint) AS wal_bytes FROM pg_ls_waldir();Also measure the WAL directory size on disk separately from the 1.17 GB whole-volume number. If the directory-size query requires privileges, run that check as the database admin or use the filesystem size. Check the WAL archive-status directory for a backlog of .ready files versus .done files.
Interpretation:
- If the WAL directory stabilizes around the configured WAL sizes while segment filenames rotate and checkpoints recycle them, there is no local cleanup fault to fix in pgBackRest.
- If it grows past the configured maximum WAL size, check for an inactive replication slot retaining old WAL, a high WAL retention setting, and unarchived .ready files / archiver status failures. The configured maximum WAL size is a soft limit and can be exceeded when WAL is retained.
- If the WAL directory is stable but the volume keeps growing, find the other directory consuming space (async spool, logs, backups, etc.). Changing WAL settings will not fix that.
Please do not manually delete files from the WAL directory; that can make the cluster unrecoverable. Share the two measurements and slot/archive status, and we can identify which of these applies before changing production configuration.
PostgreSQL docs: https://www.postgresql.org/docs/current/wal-configuration.html and https://www.postgresql.org/docs/current/runtime-config-wal.html