Deployment marked FAILED after successful build and Gunicorn boot, with no application exception
zxc22665513
HOBBYOP

a month ago

Title:

Deployment marked FAILED after successful build and Gunicorn boot, with no application exception

Hi Railway team,

I’m investigating a production deployment that was marked FAILED even though the build completed successfully and the application container appeared to start normally.

Project:

yaoyan-pos

Environment:

production

Service:

yaoyan-pos

Failed deployment:

d51f1241-f3c0-478e-b71f-a65d082ba87d

Failed deployment time:

2026-08-21 around 16:56–17:09 UTC+8

Successful recovery deployment:

18e0e211-eb65-4758-afd5-55efff5ceb92

Previous successful deployment:

3aeee8b2-6c8b-4a77-8f5d-47f86199b1ed

Builder:

Nixpacks

Runtime:

Railway V2

Start command:

gunicorn --workers 1 --timeout 180 --graceful-timeout 30 wsgi:app

Port:

0.0.0.0:8080

No explicit healthcheck path is configured.

What we observed on the FAILED deployment:

  1. Build completed successfully.
  2. Dependency installation completed successfully.
  3. Build logs explicitly show image export/push completing.
  4. Container started.
  5. Gunicorn started successfully.
  6. Gunicorn listened on 0.0.0.0:8080.
  7. Worker booted successfully.
  8. No Python/application exception was logged.
  9. No crash traceback was logged.
  10. The deployment was later terminated.
  11. Gunicorn received SIGTERM and shut down normally.
  12. Railway ultimately marked the deployment as FAILED.

During the failed deployment, our public endpoints returned 502:

/customer

/customer?entry=order

/static/customer.js

However historical Railway HTTP logs for the failed deployment return zero request records, so we cannot determine whether traffic was ever attached to the deployment instance.

The failed deployment metadata also does not expose an image digest, despite the build logs confirming that image push completed.

We immediately restored the previous known-good application commit.

Recovery deployment:

18e0e211-eb65-4758-afd5-55efff5ceb92

After recovery:

/customer → HTTP 200

/customer?entry=order → HTTP 200

/static/customer.js → HTTP 200

/daily-close → HTTP 200

The recovery used the same:

  • Railway service
  • environment
  • Nixpacks builder
  • Railway V2 runtime
  • Gunicorn start command
  • port configuration

The application change in the failed deployment was frontend-only:

  • static/customer.js
  • static/premium.css
  • templates/customer.html

There were:

  • no backend changes
  • no database migrations
  • no authentication changes
  • no order API changes
  • no stored-value/accounting changes

The failed commit also passes locally:

  • 63/63 automated tests
  • Flask app import
  • /customer template rendering
  • JavaScript syntax validation

Could you please inspect Railway’s internal deployment/platform logs for:

d51f1241-f3c0-478e-b71f-a65d082ba87d

Specifically, I would like to know:

  1. Why was this deployment marked FAILED after the container and Gunicorn successfully started?
  2. Was the deployment instance successfully registered with Railway’s proxy/network layer?
  3. Why are there no HTTP request records for this deployment despite users receiving 502 responses?
  4. Why is the image digest absent from deployment metadata even though image push completed?
  5. Was the SIGTERM initiated by Railway due to an internal readiness/network/deployment failure?
  6. Is there any Railway-internal error/event that is not visible through the CLI logs?

I have intentionally not retried the failed release because I want to determine whether this was an application issue or a Railway platform/deployment orchestration issue first.

Thank you.

Solved

1 Replies

Railway
BOT

a month ago

The failed deployment's step-by-step timeline shows the build, container creation, and app boot all completed normally, but the CONFIGURE_NETWORK step (the platform-internal step that wires traffic to the container) ran for over 7 minutes and then the deployment was marked FAILED. Your recovery deployment completed the same step in 1 second. Because network configuration never finished, traffic was never routed to the container, which is why you saw 502s, why there are no HTTP request records for that deployment, and why the SIGTERM arrived (the platform terminated the deployment after the step timed out). No healthcheck is configured on this service, so the failure is not a healthcheck timeout. It is safe to retry the failed commit, as the recovery deployment on the same service, config, and runtime succeeded immediately after.


Status changed to Awaiting User Response Railway • about 2 months ago


Railway
BOT

a month ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...