Production outage: deployment and rollback fail at CREATE_CONTAINER with no runtime logs
Anonymous
HOBBYOP

24 days ago

Production service is unavailable with HTTP 502 since September 11, 2026 at about 17:59 UTC.

Region: US West (sfo). Existing persistent volume mounted at /data.

The previous known-good deployment received SIGTERM at 17:59:07 UTC and Litestream logged a clean shutdown.

New deployment f7733f6f-c533-4dfd-afe9-7e1aecc84e73 built and published successfully. CREATE_CONTAINER started at 17:58:50 UTC, never completed, and the deployment became FAILED at 18:14:48 UTC with no application runtime logs.

Rollback deployment 5835e3f7-d37f-4829-8d15-77fff79dfc35 used the known-good image. It entered CREATE_CONTAINER at 18:14:49 UTC and became FAILED at 18:30:45 UTC, also without runtime logs.

Railway automated diagnosis completed at 18:32:41 UTC, categorized this as infra_error / Infrastructure Error, and recommended contacting support. It reports failure before startup with zero CPU/memory metrics. Its claim that the previous deployment is healthy is incorrect: direct HTTP checks return 502 and SSH cannot reach the application. Please do not rely on the dashboard Online label. Other project services show SUCCESS.

The release changed no database schema or volume configuration. A verified SQLite backup was made immediately before release. We have not restored the database, changed or recreated the volume, or started repeated deployments.

Please investigate container provisioning and the compute host associated with the volume, and help restore the known-good deployment while preserving the existing /data volume and database. This is an urgent production outage.

Additional private diagnostic details are available through an appropriate private channel.

Solved

4 Replies

Status changed to Awaiting Railway Response Railway • 24 days ago


Anonymous
HOBBYOP

24 days ago

Update: production is still returning HTTP 502, now for over an hour. At 19:05:43 UTC the owner made one deploymentRestart attempt targeting the exact known-good deployment d7afaa91-ca3c-4778-a79f-18a10bd456b8. Our recovery script did not record an acknowledgement, and we have not repeated the request. Neither deploymentEvents nor the service environmentHistory shows a new restart event; the old deployment still misleadingly shows SUCCESS/RUNNING. Please correlate that request time and deployment ID on your side. Zapier is receiving upstream events but its delivery to this application now fails with "Application failed to respond". We urgently need help recovering the service and existing /data volume without data loss. No volume changes or database restoration have been performed.


24 days ago

Your service is back up as of 19:21 UTC, running your known good image.

This was on our side. A stuck container on the machine hosting your volume never released it, so neither your new deployment nor the rollback could start. Your release wasn't the problem, and your /data volume and database are untouched.

You were also right that the automated diagnosis was wrong to call the previous deployment healthy.


Status changed to Awaiting User Response Railway • 24 days ago


Anonymous
HOBBYOP

24 days ago

Thank you, Brody. We independently confirm recovery on deployment 97e459cd-b98b-470a-82c7-eceda36ce11d: HTTP 200, authenticated CRM loads, all 336 expected files match the known-good version, and SQLite integrity_check is ok with zero foreign-key violations on the original /data volume. We did not migrate the volume or restore the production database. We are reconciling integration deliveries from the outage window. Could you confirm what intervention released the stuck container, whether a platform fix is planned to prevent recurrence, and whether any customer-side action is recommended before our next deployment?


Status changed to Awaiting Railway Response Railway • 24 days ago


24 days ago

The stuck container was forcibly cleared from the host so the volume lock could be released and a new container could start. Your release, configuration, and volume were not factors, so no customer-side action is needed before your next deployment.


Status changed to Awaiting User Response Railway • 24 days ago


Railway
BOT

17 days ago

This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!

Status changed to Solved Railway • 17 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...