Postgres Volume Randomly Deleting
lukeduell
HOBBYOP

a month ago

Hi all,

Late Thursday (~03:05 UTC Friday), the Postgres volume attached to my beta environment was deleted along with the running Postgres deployment. The deployment transitioned to REMOVED at that timestamp; the volume no longer appears in railway volume list for the environment. I did not knowingly initiate this. Confirmed:

  • No collaborators on the workspace
  • No railway volume delete invocations from any shell on my machine
  • No dashboard sessions I remember
  • No scripts running against Railway's API from my side

Environment is on the Pro plan; standard Postgres image (timescale/timescaledb:latest-pg17) with a single volume at the default data path.

Two things I'm hoping to get help with:

  1. Restore the volume within the 48h window. Docs say deletion is queued for 48 hours and an email restoration link is sent — I didn't receive one. Any way to trigger the restore server-side or resend the link? Impact is complete loss of a live trading platform's Postgres state (positions, orders, audit trail, migrations).
  2. Server-side audit info on the delete. Public API doesn't expose who/what deleted a volume. Even if this specific one can't be restored, I need to know whether the delete came from the UI, CLI, or public API, and correlate to a session, so I can prevent recurrence. Enabling PITR now per your guides, but I'd like confidence in the delete surface too.
Solved

3 Replies

Railway
BOT

a month ago

Your workspace is on the Hobby plan, not Pro. Three services in the DayTrader project were deleted around 02:58 UTC on 2026-08-26, and deleted services have no grace period and no restore path on any plan. The 48-hour delayed-deletion window applies to volumes, but because the service itself was deleted the associated volume is no longer recoverable. For the audit trail, workspace audit logs record every action with its actor (user, API token, or system), and they can also be queried via the GraphQL API. On Hobby, audit log retention is 48 hours, so given the deletion timestamp the relevant entries may have already expired. Upgrading to Pro extends audit log retention to 30 days going forward.


Status changed to Awaiting User Response Railway • about 1 month ago


lukeduell
HOBBYOP

a month ago

So what I am picking up based on information I currently have:

I queued a delete for the postgres volume on tuesday, that queued event triggers 48 hours later if done through ssh tunnel. I then manually deleted the volume in the railway UI which does NOT remove the delete from the queued deletes

Last night roughly 48 hours later the queued delete ran and deleted the volume SILENTLY. If this is the case this is awful management of notifying users of deletion. I was blind sided this morning on this issue.


Status changed to Awaiting Railway Response Railway • about 1 month ago


sam-a
EMPLOYEE

a month ago

The delete you queued Tuesday at 11:05pm ET started a 48 hour timer, and fired Thursday at 11:05pm ET. That's the deletion you found Friday. What went wrong is what happened in between. When you remounted that same volume to postgres on Wednesday afternoon, we let you do it and showed it as a healthy, live volume, but the queued delete was still armed underneath. Remounting doesn't cancel it. Only the Restore Volume link in the email does. So you ran Postgres for 33 hours on a volume we'd already scheduled for destruction, with nothing in the UI suggesting anything was pending, and no warning when it fired.

One correction on the email. It was delivered Tuesday at 11:05pm and opened at 11:29pm. That doesn't change much given everything after it told you the volume was fine, but you should know it went out.

The data is gone and we can't recover it. On your second question, the queue action gets logged but the deletion itself doesn't, so the audit log wouldn't have answered this for you even on Pro.

We're sincerely sorry. We're flagging both the remount behavior and the missing audit entry to the team that owns volumes. You lost a live database because our UI told you a volume was safe when it wasn't.


Status changed to Awaiting User Response Railway • about 1 month ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...