a month ago
Hi Railway team,
I'm seeing a reproducible data-loss issue with a persistent volume attached to my service and would appreciate help investigating.
Service: EVI-BOT (service ID 77a9d205-36f0-4b1d-8e62-a07bcf3c7dec)
Project ID: c0a5d7d4-cb0b-421a-8071-43d31e98ac55
Volume: evi-bot-volume, mounted at /app/data
Replicas: 1 (confirmed, not a multi-replica issue)
Symptom: Two specific JSON config files written to the volume at runtime (/app/data/eviGuardian/guardian-store.json and /app/data/eviShield/shield-store.json) reliably disappear after every redeploy — but only these two. A third file in a sibling directory on the exact same volume (/app/data/ticket-configs/) has survived every redeploy tonight without issue.
What I've ruled out:
Only 1 replica — confirmed in dashboard Scale settings.
No environment variable path override — verified via a runtime diagnostic log showing the actual resolved write path matches the expected volume path exactly (/app/data/eviGuardian/..., /app/data/eviShield/...).
Not a write-durability issue — writes use synchronous fs.writeFileSync+fs.renameSync (one store) and fs.promises.open+fsync+rename (the other), both verified successful (read back and confirmed immediately after write, command succeeds and reports success to the end user).
Not a timing/sync-delay issue — data written and left completely untouched for 17+ minutes survived fine; it only disappears specifically when a redeploy happens, regardless of how much time has passed since the last write.
Not a lazy-vs-eager directory creation issue — made the directories get created immediately at process boot (matching the pattern used by the files that DO survive) and it made no difference; the two files still vanished on the very next redeploy.
Reproduction: Write a file to a subdirectory under the mounted volume path while the container is running, confirm it exists on disk, then trigger a redeploy (git push via Railway's GitHub integration) — the file is gone at the very start of the new container's logs, before any app code runs.
Volume Usage metrics also show 0 B used throughout, even right after writing data that should be several KB.
Happy to provide full logs/diagnostics on request. This is affecting live paying customers, so any guidance would be very much appreciated.
Thanks,
[David]
2 Replies
a month ago
The volume mount path on your service has a trailing space character: /app/data instead of /app/data. This means the volume is mounted at a path that does not match where your application writes, so your app's writes to /app/data/eviGuardian/... and /app/data/eviShield/... land on the container's ephemeral filesystem and are discarded on every redeploy. You can fix this by editing the volume's mount path in the volume settings to remove the trailing space, then redeploying.
Status changed to Awaiting User Response Railway • about 2 months ago
a month ago
Thanks for the suggestion, but retyping the mount path fresh (selected all, deleted, retyped /app/data, saved, redeployed) didn't fix it — the two files still don't survive the redeploy. So it wasn't a trailing-space issue after all.
For reference, here's what's still happening on every redeploy: the mount path resolves correctly (/app/data, confirmed via runtime diagnostic), only 1 replica, no path override env vars, and a sibling folder on the exact same volume (ticket-configs/) persists fine across every redeploy — only two specific subdirectories (eviGuardian/, eviShield/) lose their files every time.
Could someone take a closer look at the actual storage-node level for this volume instance? Happy to provide deployment IDs or run any additional diagnostic you need.
Status changed to Awaiting Railway Response Railway • about 2 months ago
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 2 months ago