SE Asia: every deploy fails in deploy phase with no container logs since 12:23 UTC; service fully down
ccy-fls
PROOP

a month ago

Project: linchpin-prod (5567be22-1939-493a-92b3-656ae9fac597)

Service: linchpin-app (d2dac02f-37ed-4a16-b59d-75e6a2cf9eae)

Environment: production (37a35831-25a6-47bc-a169-b0630988e0c0)

Region: Southeast Asia (Metal), single volume mounted at /data

Our service has been fully down since ~12:23 UTC today (22 Aug). Since ~13:22 UTC, every deployment fails in the DEPLOY phase and the container emits ZERO runtime logs - our entrypoint never starts. This includes a plain Redeploy of an already-built image, so it is not a build issue.

Failed deployments, all with no container logs:

  • 93cf7231-c5e8-480f-a34d-6d8db3a22f2b (13:22 UTC, railway up)
  • a2d2c226-4cc6-4c30-8cf0-824f8d86d779 (13:25 UTC, railway up)
  • 0b7e1e6b (13:27 UTC, railway up)
  • a5e67415 (13:30 UTC, Redeploy of existing image)

Context: two earlier deployments today (564d14c5 at 12:23, e9d46bb7 at 12:47) DID start containers (their logs show our app's own startup check refusing - an application issue on our side, since fixed). The deployments after 13:22 are different: no logs at all, so the container appears never to be scheduled/started.

The last fully healthy deployment was 7b73d1b3 (11:08 UTC), so the platform worked for us this morning. Volume shows Status: Ready, 1.9GB/50GB. healthcheckPath /api/health with 900s timeout is configured. The public status page shows all green.

Question: is something on the platform side preventing container starts for this service or region? And can you see the actual failure reason for deployment a5e67415? We had a similar platform-side freeze yesterday (~12:50 UTC 21 Aug) that self-recovered.

This serves a school - happy to provide anything else needed.

Solved

2 Replies

Status changed to Awaiting Railway Response Railway • about 1 month ago


ccy-fls
PROOP

a month ago

Update: our signature exactly matches the resolved thread "postgres-world frozen, volume not released" (host hardware failure, us-west2, 2 days ago) — but in Southeast Asia: deploys after 13:22 UTC produce zero container logs, a Redeploy of an existing image also fails, railway ssh now says the container exited, and the volume (linchpin-app-volume, mounted /data) shows Ready while nothing can be scheduled. Please check the health of the host holding linchpin-app-volume, and if it has failed, force-release/migrate the volume to a healthy host PRESERVING ITS DATA, then tell us when one controlled redeploy is safe. Volume data integrity is the priority - it holds a school's production database.


ccy-fls
PROOP

a month ago

Resolved - root cause was on our side, not Railway's. Our deploy pipeline was shipping a build context that didn't match our own release-integrity fingerprint (a stray untracked file), so our builds were failing our own check. The dashboard's per-phase failure detail (Build > Build image) is what led us to it. Apologies for the noise, and thanks.


Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...