a month ago
Subject: Production outage — deployments stuck, three custom domains down for 2+ hours
Project: biancospino-stripe-b (25df1178-eb64-477b-bff0-9fe8928f0e8d)
Environment: production (b088ebfa-7709-4f88-9ada-95b138164681)
Service: biancospino-stripe-b (4787baf0-1626-406b-a988-7c6974bc6954)
Region: sfo, 1 replica
Edge Request ID: rNVUNEywTcCHFjCg9o6EoQ
SYMPTOM
Three custom domains in production have been returning "Application failed to respond" since roughly 00:30 UTC on 29 Aug 2026:
soggiorno.bolognarooms.com
soggiorno.borgomalvasia.com
soggiorno.villatortorelli.com
These are guest-facing pages for a hospitality business. Guests cannot open their stay pages.
STATE
Three deployments are stuck in non-final states and none is being promoted to active:
17b1f8e8-9374-46a5-b25e-867d843b6911 — created 01:35 UTC — INITIALIZING, then BUILDING
f10650ab-e613-4ee0-a48f-d496c37d357a — created 00:31 UTC — QUEUED, unchanged for over an hour
f29e2269-5dcd-4ddd-96dc-d9d4b09d3076 — created 00:04 UTC — DEPLOYING
THE APPLICATION ITSELF IS HEALTHY
Deployment f29e2269 built successfully (image pushed 00:10:26 UTC) and its container started correctly at 01:27:06 UTC. Its deploy logs show:
"INFO: Application startup complete."
"INFO: Uvicorn running on http://0.0.0.0:8080"
Background jobs were still executing normally at 01:30 UTC (scheduled tasks, outbound Stripe API calls returning 200).
So the process is up and listening on 8080, but the edge proxy has no active deployment to route to. This looks like a promotion/routing problem on your side, not an application fault. No healthcheck is configured on this service, so nothing on our side should be blocking promotion.
WHAT WE ALREADY TRIED
• Redeploy (reusing the existing build): moved a deployment out of QUEUED, but it never reached SUCCESS.
• Restart of the running deployment f29e2269: rejected with "Deployment is not restartable".
• Further redeploys only add to the queue, so we have stopped.
WHAT WE ARE ASKING
Please cancel the stale QUEUED/DEPLOYING deployments and promote a single healthy one to active so the domains resolve again.
We assume this is related to the ongoing "Deployments slow to start" incident shown on the dashboard. If so, please confirm and give an ETA. If it is not related, please tell us what on our side is preventing promotion.
Thank you.
1 Replies
a month ago
Apologies for the trouble. Your most recent deployment passed its healthcheck and is stalled at the final platform-side network-configuration step, confirming this is not an application fault. This aligns with two active incidents affecting your region: Storage and networking issues affecting some services in US West and Deployments slow to start. The stalled deployments are held by infrastructure that is currently degraded, so cancelling and promoting from our side would hit the same bottleneck. Once both incidents are resolved, trigger a single fresh deploy, which will supersede the stalled ones and restore routing to your domains.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago