14 days ago
Three same-commit deployments of a production Python/FastAPI service in asia-southeast1-eqsg3a stalled at CREATE_CONTAINER on 2026-09-21 (15:09–15:56 UTC). Image builds and pre-deploy database migrations completed successfully, but no API startup logs appeared. Metadata reported no failureStage, failureError or configErrors.
Latest attempt (UTC): created 15:49:30.464; build completed 15:49:51.910; pre-deploy completed 15:50:00.760; CREATE_CONTAINER pending from 15:50:00.900; manually aborted and REMOVED at 15:55:59.077. Earlier attempts stalled at the same step. All three attempts are now aborted.
The previous successful deployment remains online and healthy. The frontend was retained on the previous compatible version. This blocks a release; it is not a confirmed production outage.
Configuration: Dockerfile build, one replica, persistent volume at /app/storage, /health healthcheck with 100-second timeout, pre-deploy python run_migration.py, effective railway.json start command runs uvicorn app.main:app.
Root cause remains unknown: image pulling, scheduling and volume handoff need investigation. Pre-deploy success does not prove volume health because pre-deploy containers do not mount volumes.
Could Railway inspect container-creation and volume-handoff events for the selected service and advise safe recovery preserving the active deployment and files? Deployment IDs are available through a staff-only channel. Please advise before any downtime or storage changes.
3 Replies
Status changed to Awaiting Railway Response Railway • 14 days ago
Status changed to Solved t-chiab • 14 days ago
14 days ago
This issue is still unresolved. It was marked Solved by mistake. Please reopen the thread and change its status back to Awaiting Railway Response. The production deployment remains blocked at CREATE_CONTAINER; the previous deployment is still serving traffic.
Status changed to Awaiting Railway Response Railway • 14 days ago
14 days ago
The thread is reopened. All three deployments from the 15:09-15:56 UTC window show the same pattern: build and pre-deploy completed successfully, then CREATE_CONTAINER started and never progressed, with no runtime logs emitted (the container never started). Your volume is in READY state with ~93 MB used, and your previous deployment (14:20 UTC) remains online. Since all three stuck deployments are now cancelled and no build slots are held, a fresh deploy attempt would confirm whether the condition has cleared.
Status changed to Awaiting User Response Railway • 14 days ago
14 days ago
Update: the deployment issue is now resolved for us. Following your recommendation, we retried the same commit without changing the configuration, region or volume. The deployment succeeded at 16:21:08 UTC on 2026-09-21. CREATE_CONTAINER completed in approximately 144 seconds, followed by a successful healthcheck. The backend and frontend release is now live, and API health and the payment-voucher smoke test passed. We have not confirmed the root cause of the earlier stalled attempts, but the deployment blocker has cleared. Thank you for your help!
Status changed to Awaiting Railway Response Railway • 14 days ago
Status changed to Solved t-chiab • 14 days ago