a month ago
Title:
Deployment marked FAILED after successful build and Gunicorn boot, with no application exception
Hi Railway team,
I’m investigating a production deployment that was marked FAILED even though the build completed successfully and the application container appeared to start normally.
Project:
yaoyan-pos
Environment:
production
Service:
yaoyan-pos
Failed deployment:
d51f1241-f3c0-478e-b71f-a65d082ba87d
Failed deployment time:
2026-08-21 around 16:56–17:09 UTC+8
Successful recovery deployment:
18e0e211-eb65-4758-afd5-55efff5ceb92
Previous successful deployment:
3aeee8b2-6c8b-4a77-8f5d-47f86199b1ed
Builder:
Nixpacks
Runtime:
Railway V2
Start command:
gunicorn --workers 1 --timeout 180 --graceful-timeout 30 wsgi:app
Port:
0.0.0.0:8080
No explicit healthcheck path is configured.
What we observed on the FAILED deployment:
- Build completed successfully.
- Dependency installation completed successfully.
- Build logs explicitly show image export/push completing.
- Container started.
- Gunicorn started successfully.
- Gunicorn listened on 0.0.0.0:8080.
- Worker booted successfully.
- No Python/application exception was logged.
- No crash traceback was logged.
- The deployment was later terminated.
- Gunicorn received SIGTERM and shut down normally.
- Railway ultimately marked the deployment as FAILED.
During the failed deployment, our public endpoints returned 502:
/customer
/customer?entry=order
/static/customer.js
However historical Railway HTTP logs for the failed deployment return zero request records, so we cannot determine whether traffic was ever attached to the deployment instance.
The failed deployment metadata also does not expose an image digest, despite the build logs confirming that image push completed.
We immediately restored the previous known-good application commit.
Recovery deployment:
18e0e211-eb65-4758-afd5-55efff5ceb92
After recovery:
/customer → HTTP 200
/customer?entry=order → HTTP 200
/static/customer.js → HTTP 200
/daily-close → HTTP 200
The recovery used the same:
- Railway service
- environment
- Nixpacks builder
- Railway V2 runtime
- Gunicorn start command
- port configuration
The application change in the failed deployment was frontend-only:
- static/customer.js
- static/premium.css
- templates/customer.html
There were:
- no backend changes
- no database migrations
- no authentication changes
- no order API changes
- no stored-value/accounting changes
The failed commit also passes locally:
- 63/63 automated tests
- Flask app import
- /customer template rendering
- JavaScript syntax validation
Could you please inspect Railway’s internal deployment/platform logs for:
d51f1241-f3c0-478e-b71f-a65d082ba87d
Specifically, I would like to know:
- Why was this deployment marked FAILED after the container and Gunicorn successfully started?
- Was the deployment instance successfully registered with Railway’s proxy/network layer?
- Why are there no HTTP request records for this deployment despite users receiving 502 responses?
- Why is the image digest absent from deployment metadata even though image push completed?
- Was the SIGTERM initiated by Railway due to an internal readiness/network/deployment failure?
- Is there any Railway-internal error/event that is not visible through the CLI logs?
I have intentionally not retried the failed release because I want to determine whether this was an application issue or a Railway platform/deployment orchestration issue first.
Thank you.
1 Replies
a month ago
The failed deployment's step-by-step timeline shows the build, container creation, and app boot all completed normally, but the CONFIGURE_NETWORK step (the platform-internal step that wires traffic to the container) ran for over 7 minutes and then the deployment was marked FAILED. Your recovery deployment completed the same step in 1 second. Because network configuration never finished, traffic was never routed to the container, which is why you saw 502s, why there are no HTTP request records for that deployment, and why the SIGTERM arrived (the platform terminated the deployment after the step timed out). No healthcheck is configured on this service, so the failure is not a healthcheck timeout. It is safe to retry the failed commit, as the recovery deployment on the same service, config, and runtime succeeded immediately after.
Status changed to Awaiting User Response Railway • about 2 months ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago