PRODUCTION OUTAGE — Postgres volume stuck Migrating / failed US West → Southeast Asia
blackbook75
PROOP

11 days ago

Severity: P1 Production outage

Impact: Entire company cannot use our production web app (login and all pages fail). Web service is Online, but Postgres is unreachable.

Project details

Workspace: Growfox's Projects

Project: Growfox Data Analytics Tool

Project ID: 3dca3a78-180e-4227-97b8-5830ed11e932

Environment: production (61ee55c8-5a0b-4df7-bd93-7894de3c79e6)

Web service: web (43e2519d-d844-44b7-a375-5ddc0fb0b0ac) — SUCCESS / Online

Postgres service: Postgres (534f49a6-cf2f-4490-be77-47819421dd25)

Volume: postgres-volume (b73b6cb5-22f6-45b6-a975-567f0b1ad7d1) — status Migrating

TCP proxy: caboose.proxy.rlwy.net:25190

Related incident: https://status.railway.com — “Connectivity issues in US West”

What happened (timeline, GMT+7 / UTC+7)

Morning: production was working normally.

~10:10–10:15: Postgres became unhealthy (statement timeouts on django_session / deletes; then connection timeouts from web).

Redeploy attempt on US West failed with almost no deploy logs:

Deployment ID: 40e32c81-970f-4b5b-962c-9b207b34fcdf (FAILED)

We attempted region change to Southeast Asia to escape the US West incident:

Deployment ID: d7203cb0-4974-4331-82b0-0d6a764b98f0 (INITIALIZING)

UI showed “Copying to Southeast Asia…”, then:

“Failed to copy data from US West to Southeast Asia”

“An error occurred during the data migration”

Now volume remains stuck in Migrating. App cannot connect:

Internal: postgres.railway.internal:5432 → timeout / connection refused

Public proxy: TCP connects, then Postgres handshake closes unexpectedly

Why this is a platform issue

Web app config is correct: DATABASE_URL = ${{Postgres.DATABASE_URL}}

Even the previous ACTIVE/US West deployment cannot serve connections while the volume is locked in Migrating after a failed region copy.

We need the original volume unlocked and Postgres brought back online ASAP.

Region migration can wait. Restoring the previous working US West volume attachment is the priority.

Requested action (urgent)

Abort/cancel the stuck Southeast Asia volume migration for postgres-volume.

Remount/restore the original volume to the last known-good Postgres deployment on US West (or whatever region still has the intact data).

Confirm postgres accepts connections on postgres.railway.internal:5432 and caboose.proxy.rlwy.net:25190.

Please treat as production outage — our whole company depends on this app right now.

Thank you.

Solved

4 Replies

Status changed to Awaiting Railway Response Railway 11 days ago


chandrika
EMPLOYEE

11 days ago

We're sorry for the trouble. Your Postgres volume is stuck in a migrating state after the region copy from US West to Southeast Asia failed during an ongoing connectivity incident affecting US West. Your data is not lost. The volume is locked in a transitional state that we need to resolve on our side by aborting the failed migration and restoring the volume to its previous working state. We're on it. You can follow the incident here: https://status.railway.com/incident/RL8FRJE6


Status changed to Awaiting User Response Railway 11 days ago


restimo
PRO

11 days ago

it worked, then it crashed. system is unstable...


Status changed to Awaiting Railway Response Railway 11 days ago


restimo
PRO

11 days ago

does not work. crashed


chandrika
EMPLOYEE

11 days ago

Incident has been resolved now, please redeploy any services that aren't seeing recovery and if you still run into issues, please reach out to us


Status changed to Awaiting User Response Railway 11 days ago


Status changed to Solved chandrika 11 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...