a month ago
Hi,
We received the "Server Outage Affecting Your Service" notice for our Redis
service and want to follow up, because the resolution took considerably longer
than the 60 minutes stated and the failure mode was hard to detect from our side.
Affected resource
Project: (f56c073b-62ce-44b0-a5c3-d4ed4d1178cd)
Environment: production (ea755305-e8ac-4d01-8e53-4e084b3d480e)
Service: Redis (87615adf-b017-4585-ae0a-6eff919823ee)
Private IP observed by our API: 10.200.198.6:6379
Timeline (UTC)
22:33:24 Last log line from the Redis service (a routine BGSAVE). Nothing
after that — the process stopped serving entirely.~22:41 Our API begins accumulating "dial tcp 10.200.198.6:6379: i/o
timeout" on every connection attempt.22:41 - Our public checkout is effectively down: request latency goes from
00:09 ~0.3s to ~12s and card tokenization returns 503. Card payments
could not complete during this window.00:09 We mitigate on our side by removing REDIS_URL from the API service
and redeploying, so the app runs cacheless.01:14:57 Redis logs "Ready to accept connections tcp". Data was intact
(DB loaded from disk).Total service downtime: 2h41. The notice said "Expected resolution: within 60
minutes", and we received no further updates during the remaining ~1h40.
What made this hard to diagnose
During the entire outage, your control plane reported the service as healthy:
environment_status showed Redis as SUCCESS with an active deployment, and
private networking reported State: ready / Sync status: ACTIVE for
redis.railway.internal. Meanwhile no TCP connection to the service could be
established. Service metrics also kept returning data (CPU ~0, memory flat),
which reads as "idle and fine" rather than "unreachable". The only reliable
signal that something was wrong was the absence of new log lines.
Questions
-
Can you share an RCA for the hardware failure and why recovery took 2h41
against a 60-minute estimate?
-
Why were there no follow-up notifications once the initial estimate passed?
A single "still working on it" would have changed our mitigation decisions.
-
Is there a way for a service to be reported as unhealthy in the dashboard
and API when its host is down? A service showing SUCCESS while refusing all
connections is the part that cost us the most time.
-
What is your recommended posture for a Redis service with an attached
volume when its host fails — is automatic rescheduling to another host
possible, or does the volume pin the service to the failed machine? If it
pins, we would like to understand our options for reducing this blast radius.
-
Does this incident qualify for SLA credit under our current plan?
Happy to provide any additional logs or IDs you need.
Thanks,
Leo
1 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
a month ago
At 22:34 UTC on August 20, the physical server hosting your Redis had a hardware power fault and went offline instantly. Because a service with a volume runs where its data lives, it is not rescheduled automatically; recovery meant provisioning a replacement server and restoring every affected volume onto it, and that full restore is what pushed resolution past the initial estimate in the notice. Your Redis came back at 01:15 UTC with data intact, loaded from disk. We sent no further updates after the first notice, and we are sorry, you should have received one. The deployment status you saw reflects the most recent deploy result rather than live connectivity, so it does not change when a host fails; the outage notice is currently the signal for host-level incidents. To reduce the blast radius, your Redis can be converted into a Sentinel-backed high availability cluster with automatic failover from the service's Database > Config > High Availability section, currently in beta via Priority Boarding. The full guide is here.
Status changed to Awaiting User Response Railway • about 2 months ago
Status changed to Solved dizzydes90 • about 2 months ago