11 days ago
Severity: P1 Production outage
Impact: Entire company cannot use our production web app (login and all pages fail). Web service is Online, but Postgres is unreachable.
Project details
Workspace: Growfox's Projects
Project: Growfox Data Analytics Tool
Project ID: 3dca3a78-180e-4227-97b8-5830ed11e932
Environment: production (61ee55c8-5a0b-4df7-bd93-7894de3c79e6)
Web service: web (43e2519d-d844-44b7-a375-5ddc0fb0b0ac) — SUCCESS / Online
Postgres service: Postgres (534f49a6-cf2f-4490-be77-47819421dd25)
Volume: postgres-volume (b73b6cb5-22f6-45b6-a975-567f0b1ad7d1) — status Migrating
TCP proxy: caboose.proxy.rlwy.net:25190
Related incident: https://status.railway.com — “Connectivity issues in US West”
What happened (timeline, GMT+7 / UTC+7)
Morning: production was working normally.
~10:10–10:15: Postgres became unhealthy (statement timeouts on django_session / deletes; then connection timeouts from web).
Redeploy attempt on US West failed with almost no deploy logs:
Deployment ID: 40e32c81-970f-4b5b-962c-9b207b34fcdf (FAILED)
We attempted region change to Southeast Asia to escape the US West incident:
Deployment ID: d7203cb0-4974-4331-82b0-0d6a764b98f0 (INITIALIZING)
UI showed “Copying to Southeast Asia…”, then:
“Failed to copy data from US West to Southeast Asia”
“An error occurred during the data migration”
Now volume remains stuck in Migrating. App cannot connect:
Internal: postgres.railway.internal:5432 → timeout / connection refused
Public proxy: TCP connects, then Postgres handshake closes unexpectedly
Why this is a platform issue
Web app config is correct: DATABASE_URL = ${{Postgres.DATABASE_URL}}
Even the previous ACTIVE/US West deployment cannot serve connections while the volume is locked in Migrating after a failed region copy.
We need the original volume unlocked and Postgres brought back online ASAP.
Region migration can wait. Restoring the previous working US West volume attachment is the priority.
Requested action (urgent)
Abort/cancel the stuck Southeast Asia volume migration for postgres-volume.
Remount/restore the original volume to the last known-good Postgres deployment on US West (or whatever region still has the intact data).
Confirm postgres accepts connections on postgres.railway.internal:5432 and caboose.proxy.rlwy.net:25190.
Please treat as production outage — our whole company depends on this app right now.
Thank you.
4 Replies
Status changed to Awaiting Railway Response Railway • 11 days ago
11 days ago
We're sorry for the trouble. Your Postgres volume is stuck in a migrating state after the region copy from US West to Southeast Asia failed during an ongoing connectivity incident affecting US West. Your data is not lost. The volume is locked in a transitional state that we need to resolve on our side by aborting the failed migration and restoring the volume to its previous working state. We're on it. You can follow the incident here: https://status.railway.com/incident/RL8FRJE6
Status changed to Awaiting User Response Railway • 11 days ago
11 days ago
it worked, then it crashed. system is unstable...
Status changed to Awaiting Railway Response Railway • 11 days ago
11 days ago
does not work. crashed
11 days ago
Incident has been resolved now, please redeploy any services that aren't seeing recovery and if you still run into issues, please reach out to us
Status changed to Awaiting User Response Railway • 11 days ago
Status changed to Solved chandrika • 11 days ago