Orphaned container keeps serving traffic after redeploys - split-brain on volume-backed service
nyczone29-del
HOBBYOP

a month ago

Service caringplace-pickup (id 5c97fafe-a0b9-4d74-b961-f6f2f8140bef), project hearty-fascination (id 8beb5395-aaf1-4d79-b8f8-c151cf3a59ea), EU West, 1 replica, volume attached, SQLite on the volume.

After a series of rapid deploys this morning (2026-08-27 ~07:00-08:30 UTC, mix of variable changes and code pushes), an old container instance appears to still be running and receiving part of our public traffic, even though every old deployment shows REMOVED in the dashboard.

Evidence:

  • Inbound Twilio webhook POSTs to our domains are answered (correct TwiML replies reach Twilio) but never appear in the active container's logs - not even our unconditional entry log line added for debugging.
    • The same API endpoint returns DIFFERENT data via the public domains vs curl localhost:PORT inside the active container's Console (stale snapshot vs current rows).
    • ps in the active container showed the node process with ~0 CPU time while public traffic was flowing.
    • The SQLite database on the volume shows interleaved writes from two writers; a fresh connection in the Console sees all rows, while public API responses return an older subset.

Already tried, no change: multiple new deploys, Redeploy, Restart of the active deployment, enabling Teardown (0s drain), and pointing traffic at the caringplace-pickup-production.up.railway.app domain instead of the custom domains (socialenterprisetcp.org / www).

Request: please locate and terminate any orphaned/stale container(s) for this service, and confirm the edge only routes to the current deployment. This is a small nonprofit's donation-pickup system (The Caring Place, Miami) - happy to provide anything else you need.

Solved

2 Replies

Railway
BOT

a month ago

Deploys for a volume-attached service are serialized: the previous container is stopped and confirmed stopped before the new one starts, so two containers never run against the same volume simultaneously. Your deployment history confirms each prior deploy was removed before its successor launched. The edge HTTP logs also confirm that all traffic across your three domains is routing to the current deployment only, with no requests reaching any other container. There is no orphaned container to terminate. The symptoms you are seeing (divergent responses, interleaved writes) are occurring within your single running container and application layer. For hands-on help debugging that, our community at station.railway.com/questions is the right place.


Status changed to Awaiting User Response Railway • about 1 month ago


nyczone29-del
HOBBYOP

a month ago

ely on our side. Root cause: the API endpoint I was polling to verify inbound webhook writes defaults to a status filter (it lists only unmatched messages; all the "missing" rows were successfully matched and therefore filtered out). The webhook, the volume, and the routing were working correctly the whole time, exactly as your team said. There was never an orphaned container. Apologies for the noise, and thank you for the fast, accurate response - the note that all traffic was confirmed routing to the current deployment is what pushed me to re-audit our own query. This thread can be closed.


Status changed to Awaiting Railway Response Railway • about 1 month ago


Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...