a month ago
Our production service has been hard down since ~01:09 UTC today, caught in the ongoing "Deployments slow to start" incident.
Project: 7500b651-3967-4090-af4c-e3209b22d43e (kangliang-ai-assistant, environment production, US West)
Service: 812b1b6b-04e8-464c-958f-c307c21f9eda
Timeline (UTC):
- 00:42 railway up -> deployment b8af8b37 built OK, container booted healthy, then stuck at "Creating containers..." forever
-
- 01:07-01:09 platform stopped BOTH the previous healthy deployment 38516a0f (now CRASHED) and b8af8b37's container -> nothing serving, domain returns 502/timeouts
-
- 01:13 railway up again -> 2560d02b built OK, now QUEUED "Waiting for previous deployment"
-
- Tried: Restart on 38516a0f (accepted but never restarted), Abort on b8af8b37 (accepted ~01:35, still shows DEPLOYING)
This is a medical clinic's internal system (staff use it during opening hours), so downtime is painful. Could you unstick/abort b8af8b37 so the queued 2560d02b can promote, or restart 38516a0f? Volume: kangliang-ai-assistant-volume. Thanks!
3 Replies
a month ago
Apologies for the trouble. Your service is caught between two active incidents right now: Deployments slow to start, which is why the stuck deployment never completed container creation and the queued one behind it cannot promote, and Storage and networking issues affecting some services in US West, which is impacting compute in your region. Both are being actively worked on, and your service should recover as they are resolved. Follow those links for live updates.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
Update: mitigated. After the Abort of b8af8b37 finally processed, restarting the previous deployment 38516a0f worked (~01:52 UTC) and production is serving again on the old build. Remaining: 2560d02b still stuck in QUEUED with nothing ahead of it - please clear/unstick it once workers stabilize so the new build can promote. Thanks!
Status changed to Awaiting Railway Response Railway • about 1 month ago
a month ago
Glad production is back up. Your latest deployment built successfully but is stuck QUEUED behind a phantom predecessor hold with 0 of 3 concurrency slots actually in use, which is part of the same Deployments slow to start incident. It should clear as the deployment workers stabilize and work through the backlog; follow that link for live progress.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago