14 days ago
Hi Railway team,
We had a ~48-second total loss of internal network connectivity in the US region and would like help identifying the cause and reducing our exposure.
Project: lotux — c36c0c52-04e3-4333-93c0-cac52a60c586
Environment: production
Window: 2026-09-22 06:50:36Z → 06:51:24Z
What happened
Three separate services in the US region simultaneously lost TCP connectivity to our Redis database redis-us, all failing with connect ETIMEDOUT:
Service Internal source IP
lotux-background-us 10.202.12.27
lotux-worker-us 10.177.173.172
lotux-ws-us 10.150.197.44
All three connect to redis-us at 10.213.175.174:6379. Connectivity restored on its own at 06:51:24Z. No deploys, restarts, or config changes were made by us before, during, or after.
What we have already ruled out
redis-us was healthy throughout. Its own logs show uninterrupted BGSAVE snapshots every 60 seconds across the entire window (1542 keys). The Redis process never restarted and never stopped serving.
DNS was clean. railway logs --dns shows 56 queries in that window, 100% NOERROR.
Your network flow logs show no drops during the outage. railway logs --network --filter "@dropped:true" returns nothing between 06:50:36Z and 06:51:14Z. The only drops are 4 NO_SOCKET ingress packets at 06:51:15–17Z, which are a consequence of our client having already torn its sockets down — not a cause.
Not Redis-specific. During the same window, lotux-worker-us also failed to reach our EU Redis (connect ETIMEDOUT). That suggests the problem was on the US egress side generally, not on the path to redis-us.
Regionally isolated. Our EU (lotux-api, lotux-background-eu) and SG services logged zero errors in the same window.
Not on your status page. status.railway.com reports 100% September uptime for US East, US West, and "Networking — Private".
So from our side: TCP SYNs went out and nothing came back, with no DNS failure and no packet-drop record.
A related, smaller event
On 2026-09-16 17:43:30Z our EU service lotux-ws saw the same connect ETIMEDOUT signature for roughly 5 seconds before self-recovering. Two events in six days of retained logs.
Our questions
Was there a known network event in the US region during 2026-09-22 06:50:36Z–06:51:24Z? If it was below your public status threshold, we would still appreciate confirmation that something occurred.
Region placement: our project metadata shows both us-east4-eqdc4a and us-west2. Could you confirm which region each of lotux-ws-us, lotux-worker-us, lotux-background-us, and redis-us is actually deployed in? If any of those services sit in a different region from redis-us, we would like to know — we assumed they were co-located.
Is a ~48-second private-network interruption within expected behavior for our plan, or does it indicate a problem with our specific placement or host?
What can we do on our side to reduce exposure? Specifically: would moving these services to the other US region help, does adding replicas provide any protection against this failure class, and are there connection-level settings you recommend for services that hold long-lived Redis connections?
Is there a way for us to be notified when this class of event happens, rather than discovering it from our own application logs?
Context: we run a trading-execution platform. We are pre-launch with no customer traffic in the US region, so this event caused no user impact — which is exactly why we want to understand it now, before it can. We are trying to decide whether to keep building on the US region.
Happy to provide full application logs, deployment IDs, or anything else useful.
Thanks,
Minh — LOTUX
1 Replies
14 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 14 days ago
9 days ago
Hi Minh,
Here is a breakdown of what caused this ~48-second outage, why your metrics looked clean, and concrete steps to safeguard your trading platform.
- Root Cause Analysis: Why 48s, Clean DNS, and Zero Drops?
The 48-Second Signature is Linux TCP SYN Retries: When an application attempts a new TCP handshake during an unresponsive network window, Linux does not fail immediately. It retries with exponential backoff: 1s + 2s + 4s + 8s + 16s = ~31s, reaching connect ETIMEDOUT at ~48–60 seconds depending on kernel tcp_syn_retries (typically 5 or 6). The underlying network blip likely lasted only 15–25 seconds, but active and queued connection attempts backed up and failed at the 48s mark.
Why Flow Logs showed no dropped packets: Railway's private network runs an eBPF/WireGuard mesh overlay across physical hypervisors. When an egress route or WireGuard peer tunnel between hypervisors suffers a transient route flap or packet loss on the underlying transit ISP, the local interface still accepts and transmits the packet. The packet is dropped in-flight across the mesh peer path, not rejected at the local container veth interface — which is why @dropped:true showed nothing.
Cross-Region / Cross-Host Placement: Your metadata lists both us-east4-eqdc4a (Equinix East Coast) and us-west2. If redis-us was provisioned in us-east4 while your worker/ws services landed on us-west2, traffic is traversing thousands of miles over an encrypted cross-region tunnel. Furthermore, because lotux-worker-us simultaneously failed to reach the EU Redis, the node hosting your US services experienced a transient egress route reconvergence or host NIC flap.
- Answers to Your Questions
Q1 & Q3: Was it a known event / Is 48s acceptable?
48s of total packet loss is not normal steady-state behavior, but it matches a known pattern: hypervisor-level BGP failover or WireGuard peer re-key/route re-convergence on the host machine. Because it was an intermittent routing brownout on a specific host rather than an entire datacenter outage, it did not trip the public status page threshold.
Q2: Determining Actual Service Placement
Railway does not always pin all services of a project to the same host or availability zone unless explicitly configured. You can check the exact physical datacenter for each container by running:
bash
railway run --service env | grep -E "RAILWAY_REGION|RAILWAY_ZONE"
Or check the container's latency to Redis using ping 10.213.175.174 from inside the service:
< 1.5ms: Co-located in the same datacenter/rack.
50–70ms: One service is in us-east4 and the other is in us-west2.
Q4: What you can do to eliminate exposure
Strict Region Co-location: Ensure lotux-background-us, lotux-worker-us, lotux-ws-us, and redis-us are explicitly deployed in the same region (e.g., us-east4-eqdc4a). In your service settings, lock the region per service rather than allowing automatic multi-region distribution.
Connection-Level Redis Settings (Critical for Trading Platforms): Tune your connection pooling and retry logic to recover immediately without waiting for the 48-second OS timeout:
connectTimeout: Set to 5000ms (5s). Do not let it hang for the default OS TCP timeout (~48s).
keepAlive: Set TCP KeepAlive to 10000ms (10s) so dead sockets are detected and recycled before you try to write to them.
maxRetriesPerRequest: Set a finite number (e.g., 3–5) with an exponential backoff jitter:
ts
// Example for ioredis
const redis = new Redis({
host: 'redis-us.railway.internal',
port: 6379,
connectTimeout: 5000,
keepAlive: 10000,
retryStrategy(times) {
return Math.min(times * 100, 3000); // retry after 100ms, 200ms, etc.},
reconnectOnError(err) {
const targetError = 'READONLY';
if (err.message.includes(targetError)) return true;
return 1; // reconnect on transient error}
});
Fallback & Circuit Breaker: Implement an in-memory buffer or circuit breaker for non-critical jobs so a 15-second network flap doesn't crash the worker lifecycle.
Q5: Real-time Alerting
Instead of waiting for application crashes, set up an active Synthetics / Healthcheck probe: