Hard-down: deployment stuck in 'Performing healthchecks' 50+ min; container up and listening; predecessor stopped
joosebox
HOBBYOP

4 hours ago

Related to incident 8GL2R2U5 (deployments slow to start). Our production service is HARD-DOWN and stuck in a state only Railway can clear.

Project: romantic-youth (e048dbf1-9793-4f39-b1f8-74f8735348f6)

Service: recireel-recovery (3ff16808-8ddb-468c-a15a-92dc5ab18ef8), environment production, US West, 1 replica

Stuck deployment: 53b7371e-86c2-4a15-a533-53bf5e888d55 (redeploy via CLI at 00:33:53 UTC)

Timeline (UTC, Aug 29):

  • 00:33 redeploy created; progressed to 'Performing healthchecks...' and has been stuck there 50+ minutes.
    • 00:55:53 the cutover machinery SIGTERM'd the ACTIVE deployment's container (deployment 7eb79167-1357-4aa1-9ec9-0753e9e4eed6) - it drained cleanly and stopped ('Stopping Container'). Since then NOTHING serves our domains: both api.recireel.com and recireel-recovery-production.up.railway.app time out (TCP connect OK, zero response bytes).
    • Meanwhile the NEW container DID boot successfully: volume mounted 01:05:05, boot store guard all green, 'Server listening' at 01:11:43, internal /health responding. The volume I/O degradation we saw earlier (SQLite commits 60,000ms+) recovered at ~01:14 - commits are back to <100ms.

So: healthy listening container + stopped predecessor + deployment state machine wedged at healthchecks = production outage only your side can clear.

Request: please force-complete/promote deployment 53b7371e (or re-run its healthcheck phase) so ingress points at the already-healthy container, or advise the fastest safe path. Consumer iOS app backend; customer-visible degraded since ~00:23, hard-down since 00:55.

Awaiting User Response

6 Replies

Railway
BOT

4 hours ago

We're sorry for the trouble. There is an active incident, Deployments slow to start, affecting deployments across US West and other regions due to deployment worker instability. Your deployment's build, container creation, and healthcheck steps all completed, but promotion to active has not progressed, which is consistent with this incident. We are working to stabilize the system and clear the backlog. Please follow that link for live updates.


Status changed to Awaiting User Response Railway about 4 hours ago


joosebox
HOBBYOP

4 hours ago

Update 01:35 UTC: RESOLVED for us. Traffic restored at 01:25:54 UTC - once your networking fix propagated, probes reached the 53b7371e container, and ingress flipped to it (stable external 200s verified). We aborted our queued extra redeploy (bc9ce8c7) to avoid another cutover mid-incident. Remaining cleanup on your side: the service still shows 53b7371e as 'Deploying' even though its container is serving - please tidy when the backlog clears. Customer-visible impact: ~00:23-01:25 UTC.


Status changed to Solved Railway about 4 hours ago


Status changed to Awaiting Railway Response Railway about 3 hours ago


Railway
BOT

3 hours ago

Apologies for the trouble. The Deployments slow to start incident caused the delays across your deployment's steps, and your stuck deployment has now fully promoted to SUCCESS with networking configured. Separately, there is a new incident affecting US West: Storage and networking issues affecting some services in US West, which may be worth monitoring given your region.


Status changed to Awaiting User Response Railway about 3 hours ago


joosebox
HOBBYOP

3 hours ago

URGENT REGRESSION 01:47 UTC - NOT resolved, and now data-availability-threatening. When 53b7371e was promoted to active (~01:34 UTC) the public ingress path BROKE AGAIN: api.recireel.com and the service domain time out (TCP connect OK, zero bytes) while your internal health checkers still reach the container. Worse, volume I/O has collapsed on container 370df75540e2 - store SQLite commit durations 01:33-01:36 UTC: 16,643ms / 5,143ms / 40,378ms / 95,324ms / 85,851ms. 95-second disk commits on a singleton production volume is a data-availability risk, not just routing. We are deliberately NOT redeploying (removed our own queued redeploys) to avoid a volume detach/attach while your storage/network layer is unstable. Please: (1) fix/reroute ingress to the promoted container; (2) check the host/storage backing recireel-recovery-volume. Hard-down for customers since ~01:35 UTC - second hard-down window tonight.


Status changed to Awaiting Railway Response Railway about 3 hours ago


Status changed to Solved Railway about 3 hours ago


joosebox
HOBBYOP

3 hours ago

Recovery confirmed 02:01:45 UTC after your storage-infrastructure fix: 10/10 external 200s, store commits back to ~60ms (from 95,324ms), same container, no restore needed. Customer-visible windows tonight: ~00:23-01:25 and ~01:35-02:01 UTC. Solved stands. Thank you.


Status changed to Solved Railway about 3 hours ago


Status changed to Awaiting Railway Response Railway about 3 hours ago


Railway
BOT

3 hours ago

Apologies for the disruption tonight. The first window (00:23-01:25 UTC) was caused by the Deployments slow to start incident, and the second (01:35-02:01 UTC) by a storage and networking issue in US West that has since been resolved. Glad the service and volume I/O are fully back to normal with no data loss.


Status changed to Awaiting User Response Railway about 3 hours ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...