Volume stuck unattachable - every volume mutation in one project is silently dropped; need to recover the data
rafalves90
HOBBYOP

2 months ago

TL;DR: in project vivacious-nurturing (21fbd9bd-f5d3-43dc-bf22-d655adf2baed, env production f9a88663-a207-42bb-a716-33d9a6a96492, us-west) every volume mutation is accepted and then silently dropped. That made the attached volume unmountable and every deploy hang at CREATE_CONTAINER. I already migrated production to a brand-new project (where the exact same operations worked on the first try), so uptime is no longer the issue - I just need the data out of the old volume: worker-volume, a48eb6b5-b07d-4cef-baff-a0bab4b4c940, 105 MB, mount path /data, status Ready, currently 'Attached to: N/A'.

TIMELINE

Until Aug 16 the service Worker (5d8c926b-80f0-4556-9974-30cc8e881139) ran fine for two weeks with this volume mounted at /data. From Aug 16 ~23:49 UTC, EVERY deployment: build succeeds, deploy hangs at CREATE_CONTAINER, zero runtime logs, FAILED after ~15-20 min. 8 consecutive failures. Your in-dashboard Diagnosis said: 'the container fails to boot regardless of whether Railpack or a custom Dockerfile is used, so the root cause is likely in how the container starts rather than the build itself.'

WHAT I RULED OUT, ONE VARIABLE AT A TIME

  • Code: a minimal python:3.13-slim Dockerfile never emitted a single log line either.
    • Builder: failed with Railpack 0.36.4 and with my own Dockerfile (builder DOCKERFILE). Possibly a separate bug: neither the RAILPACK_VERSION service variable nor build.railpackVersion in railway.json changed the build driver - build logs kept saying 'using build driver railpack-v0.36.4'.
    • Start command: failed with a custom start command and with the field empty (image CMD).
    • Source: failed from GitHub and via 'railway up' (local snapshot).
    • Billing: workspace was on an expired trial; upgraded to Hobby - still failed.
    • Canary: a new service from a public nginx image deployed SUCCESS in seconds.
    • Same code, no volume: a new service worker2 from the same repo/Dockerfile boots and runs (it only crashes because /data is missing - expected).
    • Detaching the volume: applied in ~10 seconds, and Worker then went Online instantly - first success since Aug 16.

THE ACTUAL DEFECT: VOLUME MUTATIONS SILENTLY DROPPED IN THIS PROJECT

  • Attach volume to worker2 via dashboard: stuck at 'Changes are being applied...' for 20+ min. The state survives a hard reload and the modal offers no cancel/discard.
    • Attach via CLI: prints 'Volume worker-volume attached to service worker2', but 'railway volume list' still shows 'Attached to: N/A' and no deployment is ever created.
    • Same CLI attach against a third service (canario-nginx, which is Online): same false success, same no-op.
    • Creating a BRAND-NEW empty volume in this project also vanished: CLI printed 'Volume worker2-volume created for service worker2 ... at mount path /data' and it never appeared in 'railway volume list'.
    • An earlier staged change on Worker had failed with 'An unknown error occurred'.
    • 'railway volume files list' refuses because the volume is not attached, so I cannot read the data that way.
    • Control test: I created a NEW project, added a service from the same repo and a new volume - created, attached and mounted on the first attempt, container booted normally. So this is scoped to the old project, not to my account, my code, or my Dockerfile.

WHAT I NEED

Recover the contents of volume a48eb6b5-b07d-4cef-baff-a0bab4b4c940 (a SQLite database under /data). I understand from another thread today that there is no platform-side export and that the only supported way to read a volume is to mount it on a service - which is exactly the operation that is broken here. So could someone please unwedge whatever transaction is stuck on that project/volume, so that I can attach it to a service again and copy the data out?

Nothing in vivacious-nurturing is in use anymore (production has moved), so the project can stay exactly as it is for your investigation. I have not deleted anything - please do not delete the volume.

Solved

6 Replies

Status changed to Awaiting Railway Response Railway • about 2 months ago


dizzydes90
EMPLOYEE

2 months ago

Good news first: your data is safe and reachable right now. worker-volume (about 105 MB, your SQLite database) is intact and healthy, and it's currently mounted at /data on your canario-nginx service, which is deployed and running. You can open a shell into canario-nginx with railway ssh and copy the database out of /data. This was a platform-side issue with volume operations in that project, not anything in your code or config, and it has since cleared for your service.

One caution: if you've got any uncommitted staged changes in this environment, one would mount worker-volume onto worker2. Don't apply that one, or it'll pull the volume off the service you're reading from. Copy your data out from canario-nginx first.

The Railway Team


Status changed to Awaiting User Response Railway • about 2 months ago


rafalves90
HOBBYOP

2 months ago

Thanks - confirmed. The volume was mounted on canario-nginx, I copied the SQLite database out with railway ssh, and the data is fully recovered. I did NOT apply the pending worker2 mount. Much appreciated.

Reporting back on one thing, in case it helps: the same CREATE_CONTAINER signature came back TODAY in a DIFFERENT project - the new one I migrated to. I do see the 'Deployments are slow to progress' incident banner now, so this may simply be that incident; if so, ignore the rest.

Project odds-ev-prod (cc4b4e5e-69c5-425e-8e51-d9e2eae824fe), env production (f51b7827-1720-44cd-baaa-9b16db4ee7b5), service worker (549669fe-9ad1-4ca2-9411-6b1806c66edc), us-west, one volume at /data. It deployed fine on Aug 17 (f1fc4ef9) and has been running since. Today a config-only push failed twice, identical signature: build succeeds and the image is pushed, then it hangs at CREATE_CONTAINER with zero runtime logs and goes FAILED after ~15-20 min. Failed deployments: 9d97d4b1-2847-4147-aacb-31c9e01cbf8c and f883541e-0d7a-4607-9f07-49e8fcbec73a. Same Dockerfile, same start command, tests green.

Good news this time: the running container was NOT torn down, so the service stayed up throughout.

Question: is today's failure the ongoing deployments incident, or the same volume-path issue as before? If it is the volume path, the common factor across both projects is 'redeploy of a service that already has a volume mounted'. I will leave the two failed deployments in place in case they are useful to trace.


Status changed to Awaiting Railway Response Railway • about 2 months ago


sam-a
EMPLOYEE

2 months ago

Glad the data is fully recovered. Today's failures in the new project are the active deployment incident, not the volume bug. The original issue was stuck container removals blocking every commit workflow in that project. Today's deployments are stalling earlier in the pipeline (during build/image stages), which matches the Deployments are slow to progress incident affecting US West. Your running container was not torn down because the failure happened before the deploy reached the container-swap stage. Once that incident resolves, your deploys should progress normally.


Status changed to Awaiting User Response Railway • about 2 months ago


Status changed to Solved sam-a • about 2 months ago


rafalves90
HOBBYOP

a month ago

The container-removal bug struck AGAIN — third project affected, second forced migration in 6 days.

Timeline of this recurrence (project odds-ev-prod, cc4b4e5e-69c5-425e-8e51-d9e2eae824fe, service worker 549669fe):

• Aug 23 ~20:20 UTC: a routine deploy logged "Stopping Container", the running container was killed, and no new container ever started.

• 3 consecutive deployments FAILED with build OK and zero container logs (deployments 1b92bab5, 1bdbfab8, 03eaee4c) — the exact "stuck container removals" failure mode sam-a described above for our first dead project (vivacious-nurturing).

• status.railway.com showed "Fully Operational" the whole time (unlike the Aug 18 incident, which did NOT kill the running container).

• The same commit/Dockerfile deploys fine elsewhere: we migrated to a fresh project (odds-ev-prod2, dada9aed) and the identical image boots there (4 successful deploys today).

Two asks:

  1. Root cause: this is now the 2nd project killed by stuck container removals in a week (vivacious-nurturing Aug 16 → migrated Aug 17 to odds-ev-prod → killed Aug 23). Migrating projects every few days is not sustainable — is there a fix or a flag for our account/workload?

  2. Data recovery: volume worker-volume in odds-ev-prod is now detached and holds ~6h of Sunday data we need. Last time (Aug 18) you unstuck the volume and we recovered it by attaching to a live service. Please unstick this one the same way. We will not delete anything in that project until then.


Status changed to Awaiting Railway Response Railway • about 1 month ago


sam-a
EMPLOYEE

a month ago

Your data is reachable right now. worker-volume is mounted on worker-b, which is up, so railway ssh into worker-b and copy the SQLite database out of /data. Don't apply the pending change sitting in that environment, it would move the volume onto worker and pull it off the service you're reading from.

Same fault as before: a container from your Aug 19 deploy couldn't be stopped, so its deployment stayed marked as running and every new deploy spent about 16 minutes trying to swap it out. That stuck deployment has been cleared, and worker can deploy again.

On prevention there's no good answer to give you. There's no account or workload flag to set, and it isn't tied to your code or your project, your two cases landed on different machines, which is why migrating didn't protect you. A recurrence can be cleared from our side in minutes once the failed deployment IDs are known, so rebuilding a project isn't the remedy.


Status changed to Awaiting User Response Railway • about 1 month ago


rafalves90
HOBBYOP

a month ago

Confirmed: data fully recovered. I pulled both the SQLite database and the raw payload archive out of /data on worker-b via railway ssh.

One note for your telemetry: when the volume was mounted onto worker-b, that service (built from our production repo) booted a full second instance of our worker and spent ~20h polling providers and pushing alerts against the old database. No harm done on our side, and I have deleted worker-b now that the data is out.

Thanks for the quick unstick and for the prevention guidance. On any recurrence we will post the failed deployment IDs here immediately instead of migrating projects. Marking this solved.


Status changed to Awaiting Railway Response Railway • about 1 month ago


Status changed to Solved rafalves90 • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...