a month ago
Service: american-golf-archive / aga-runtime / production
Region: US East
Topology: one replica with persistent volume aga-runtime-volume
Current safe state
- Active deployment b0ad593c-b28a-42af-8cc2-8bdbcb6c7b87 is SUCCESS and returns HTTP 200 at /healthz after a manual Restart.
- Active image digest: sha256:931e06da729d818850b087dfbfd280b99889a6fec9b3a1bfb5bcbf0b9743413e.
- The mounted volume contains durable application state and must not be detached, reset, replaced, or deleted.
Problem
We need one controlled same-image deployment so the running container receives updated environment variables. The control plane has deleted LOGOPIPE_NEWSLETTER_POSTAL_ADDRESS and updated GIT_COMMIT_SHA, but restarting the old deployment correctly preserves that deployment's old environment snapshot. Two non-overlapping redeployments of the exact same successful image failed before application startup:
- 469c2947-9f57-43ff-9302-163981f0d71b
- 8b488794-7eca-47c6-8083-4b8b5c76a8e4
For both attempts:
- Build/image publication completed.
- Deployment entered the volume/container handoff.
- Runtime/deploy logs repeated only "Mounting volume on: .../vol_5ubjbnusjmg46ykz" and "Starting Container".
- No application startup logs appeared.
- The public service stopped responding while the new deployment held the singleton volume.
- Network healthcheck failed after approximately 4 minutes 52 seconds.
- After the failed attempt was stopped, restarting the prior active deployment immediately restored /healthz HTTP 200.
The second attempt was started only after the first was terminal and the service had been restored. No overlapping deployment remained. It reproduced the same failure.
Requested help
Please investigate the volume attachment/migration state for aga-runtime-volume and explain or correct why a same-image redeploy cannot mount the existing volume and start. We need a safe path for exactly one deployment of the existing image with current control-plane variables, preserving every byte on the volume and avoiding another extended outage.
Please do not detach, delete, reset, or replace the volume. The service is healthy on b0ad593c now. If a platform-side action is needed, please coordinate it before changing the active deployment.
1 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 2 months ago
a month ago
Going off your own log excerpt, I don't think the volume is what's failing here.
Mounting volume on: .../vol_5ubjbnusjmg46ykz followed by Starting Container is the mount succeeding. A stuck attach, or a volume still held by another container, doesn't get as far as printing the mount line, let alone starting the container. And if those two lines genuinely repeat in the log (your wording is ambiguous on that point), then the container is being restarted in a loop and the process is exiting immediately after start. Either way what you have is a boot that produces no output, not a volume handoff that never completes.
The healthcheck failure at ~4m52s is downstream of that. Nothing ever listened on the port, so the check simply ran out its window, which lines up with the 300s default. That's the symptom that got surfaced, not the cause.
Which leaves the difference you already identified yourself. A Restart replays that deployment's stored environment snapshot; a redeploy composes a fresh one from current control-plane state. You've said what changed there: LOGOPIPE_NEWSLETTER_POSTAL_ADDRESS deleted, GIT_COMMIT_SHA updated. Same image, same digest, same volume, same region, different env. If the app reads that variable at import or boot time and hard-fails when it's absent, you get precisely this shape: immediate exit, no application logs, healthcheck timeout, and a perfectly healthy older deployment the moment you restart it onto the old snapshot.
The reason the traceback isn't in your logs is that stdout is block-buffered when it isn't a TTY. A process that dies during boot never flushes, so whatever it printed is discarded and you're left with only the platform's own lines. Fix depends on the runtime:
- Python:
PYTHONUNBUFFERED=1, or invoke withpython -u -
- Node: usually fine, but a
process.exit()inside a config validator can still cut off a pending stdout write
- Node: usually fine, but a
-
- generic: send boot errors to stderr, or wrap the start command as
sh -c 'exec your-app 2>&1'
- generic: send boot errors to stderr, or wrap the start command as
Also worth ruling out: if the app logs to a file on the volume rather than stdout, the crash is sitting in that file and Railway will never surface it. Worth an ls -la over the log directory next time you're in a working container.
A test that costs nothing and touches no bytes on the volume:
- Put
LOGOPIPE_NEWSLETTER_POSTAL_ADDRESSback with its old value, or an empty string if the code only checks presence. - In the same edit add
PYTHONUNBUFFERED=1or the equivalent, so that if it still fails you get the real error this time instead of silence. - Redeploy once.
If it boots, that's your answer, and you've also had the single controlled deployment with current variables that you asked for. Then remove the dependency in code, ship that, and only afterwards delete the variable. That ordering is what avoids a second outage.
One expectation worth setting: a service with a volume can't run two deployments at once, so the old container has to be stopped before the new one can mount. The public downtime during each failed attempt is inherent to that rather than a sign of something broken, which is also why it's worth getting the boot right in one shot instead of iterating through redeploys.