a month ago
Subject: Public proxy to Redis service stopped passing connections while the service stayed healthy (~50 min, 2026-09-02)
Window: 2026-09-02, approximately 00:25–01:16 UTC. Restored ~01:16 UTC by a manual restart of the Redis service.
What happened:
- Our app (same project/environment) began timing out on every Redis operation; its logs show repeated "Redis unavailable" errors throughout the window.
- The Redis service itself was healthy the entire time: its logs show the normal periodic background save cycle continuing on schedule with zero errors.
- A manual restart of the Redis service restored connectivity immediately, with no change on our side — consistent with the TCP proxy / edge binding in front of the service having gone stale, and the restart forcing it to re-register.
Questions:
- Can you confirm from your side whether the TCP proxy / edge binding for this service stalled during that window, and what caused it?
- Is this a known issue with TCP proxying, and is a fix planned?
- Is there a mitigation on our side short of restarting the service (which also disturbs the evidence)?
Context: we are pre-launch on Railway (production runs in the same project) and have added our own hourly connection probe on this exact path; we will append dated timestamps here if it recurs.
1 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 26 days ago
25 days ago
Not Railway staff, just another Pro user running production here, so here is what I can and cannot tell you.
1 & 2 — confirming the stall / known issue. Only staff can look at the proxy side. To make that possible, add to this thread the service id, the region the Redis service runs in and the exact xxx.proxy.rlwy.net:port endpoint. What you describe (Redis healthy, zero errors in its logs, every operation timing out, immediate recovery after a restart with no change on your side) matches the failure mode where the TCP proxy keeps a stale upstream mapping for a deployment instance until a new instance registers (restart or redeploy). There have been platform incidents of that kind before; whether that is what hit your service on 2026-09-02 is something only they can confirm from the proxy logs. Keep in mind the status page only lists widespread incidents, so "nothing on the status page" does not rule it out.
3 — mitigation without restarting. The big one: if the app runs in the same project/environment, do not go through the public TCP proxy at all. Use the private network: redis.railway.internal:6379. The Redis template already exposes it as REDIS_URL; the proxy endpoint is REDIS_PUBLIC_URL. Private traffic never touches the proxy, has lower latency and is not billed as egress. Two gotchas: .railway.internal names resolve to IPv6 only, so ioredis needs family=0 in the connection string (redis://default:pass@redis.railway.internal:6379?family=0); and the private network takes a moment to come up when a container starts, so keep a retry on the first connect. Keep the public proxy for external access only (your laptop, redis-cli, a dashboard).
Also worth doing:
- Make your hourly probe test both paths (private and public). If private answers and public does not, you have isolated the proxy without touching the service, and you can post that here as evidence.
- Client side: set a connect timeout and a bounded reconnect strategy so a stalled path fails fast instead of every operation hanging.
- If you ever need to "kick" the proxy again, a redeploy of the Redis service does the same as a restart (new instance, proxy re-registers); there is nothing gentler on the user side today. Before doing it, capture
nc -vz <proxy host> <port>andredis-cli -h <proxy host> -p <port> --latency: a TCP handshake that completes but never gets a PONG points at proxy → upstream, a refused connection points at the proxy itself.
One thing to rule out on your side before assuming the proxy: compare INFO clients (connected_clients) during a probe failure with a normal value. A connection leak in the app pool (half-open connections piling up through the proxy) also looks like "Redis fine, everything times out", and a Redis restart clears that too. If restarting the app would have fixed it as well, the proxy was not the culprit.