2 months ago
CRITICAL CASE — PRODUCTION SITE OFFLINE FOR CUSTOMERS
Project: accomplished-tenderness
Environment: production
Affected service: Auto (ID: 0eb5bf1f-51e9-42d1-adb1-afdb74b8a905)
Impact window: since 2026-08-20 22:53 UTC
=== SITUATION ===
My production site is completely offline. Customers cannot access it.
Railway notified me of a hardware failure affecting my service at 22:53 UTC.
The dashboard status shows "Online" but the site does not load and every
restart or deploy attempt fails.
=== SPECIFIC ERROR ===
Any restart attempt returns:
"stacker production-stacker-cssv9a-metal-052-02 (10.66.1.2) unavailable:
rpc error: code = Unavailable desc = connection error: desc =
"transport: Error while dialing: dial tcp 10.66.1.2:50052: i/o timeout"
The infrastructure machine 10.66.1.2 is not responding.
=== WHAT I HAVE ALREADY TRIED ===
- Restart via dashboard → fails with timeout error on machine 10.66.1.2
- Deploy new commit → same timeout error
- Deploy rollback to previous commit (068d4d11) → in progress, but may
also fail for the same reason
=== WHAT I NEED ===
- Investigate why machine 10.66.1.2 is still offline AFTER the
completed hardware maintenance
- Migrate my service to another healthy infrastructure machine
- Restore my service to an online state as quickly as possible
=== IMPACT ===
Production site offline → customers cannot access
Critical time: already offline for ~3 hours
Please prioritize as CRITICAL URGENCY.
3 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
UPDATE - 3+ hours offline, critical impact
My production site remains offline since 22:53 UTC. The rollback deployment
also failed in CREATE_CONTAINER phase at 00:03:15 UTC and is still pending
after 6 minutes with zero error logs.
EVIDENCE THIS IS NOT ISOLATED:
-
User foreveryoung2149: Shopify app also down (thread: why-is-my-shopify-app-down-59449a6b)
-
User aatreviso: Postgres frozen, volume stuck (thread: postgres-world-frozen-volume-not-releas-17f88edc)
-
User girlssailing: Postgres stuck after hardware-failure marked resolved, 7 days down
(thread: postgres-stuck-exited-after-hardware-f-323363d1)
This pattern suggests infrastructure-wide recovery issues from the hardware failure,
not isolated service problems. Multiple users affected, all unable to restart/deploy,
all with CREATE_CONTAINER hanging indefinitely.
CURRENT BLOCKER:
-
Infrastructure machine 10.66.1.2 still returns connection timeouts
-
Error: "stacker production-stacker-cssv9a-metal-052-02 (10.66.1.2) unavailable:
rpc error: code = Unavailable desc = connection error"
-
This machine should have been recovered during the 22:53-23:53 UTC maintenance window
REQUEST:
- Immediately investigate machine 10.66.1.2 status
- Migrate my service to healthy infrastructure
- Allow my deployment to progress past CREATE_CONTAINER
- Review recovery status of other affected services
My business is losing revenue with every minute the site stays down.
Escalate as CRITICAL PRIORITY.
2 months ago
Unfortunately we're having an incident with the host you're project is currently on and fixing it is going to take another hour potentially.
Please accept my apologies we have several platform engineers investigating this right now. Will update here as I know more.
Status changed to Awaiting User Response Railway • about 2 months ago
a month ago
Your Auto service is back up and has been running since it redeployed successfully at 01:04 UTC. The volume shows your data intact, about 3.4 GB, and it has been taking new writes through the morning, so your site should be serving customers normally again. Sorry for how long that outage ran. It's worth double checking your most recent records to confirm nothing from the outage window is missing.
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago