A live customer application facing connectivity issues at peak hours in Jakarta
Anonymous
HOBBYOP

4 days ago

Subject: Edge proxy stalling TCP/TLS handshakes (10–38s) in Southeast Asia, causing errors for most requests to a healthy service at peak

Summary

Since about 2026-09-30, during peak hours in Jakarta (around 04:00–06:00 UTC), a large share of requests to our service fail at Railway's edge. Our app is healthy and fast. The time is being lost on the TCP connection and TLS handshake at the edge, before requests reach our container. We haven't deployed or changed config in two weeks.

Evidence

  1. Railway HTTP logs show most requests dropped at the edge. For 2026-10-01 04:30–04:45 UTC we sampled 5,000 HTTP log lines: 1,653 were 200 and 3,346 (about 67%) were 499. 92% of the 499s show 1ms latency. The same window on earlier days:

    • Sept 28: 1 x 499 out of 5,000
    • Sept 29: 0
    • Sept 30: 12
  2. The app is healthy. railway metrics for production over the last hour:

    • CPU peak 0.02 of 8 vCPU; memory peak 19 MB of 8 GB
    • Latency for requests that reach the app: p50 1ms, p95 3ms, p99 19ms
    • No restarts

    App logs in the same window show median 14ms and p95 59ms. The app's only errors are context canceled / pq: canceling statement due to user request, meaning the caller had already disconnected.

  3. Direct connections to the edge stall on TCP and TLS, on two different services and edge IPs. Tested at [time] UTC on 2026-10-01 with curl, bypassing Cloudflare:

▎ | Target | TCP connect | TLS done | Result |

▎ |-------------------------------------------------|---------------|------------|----------------------------------------|

▎ | Prod custom domain (Cloudflare bypassed) | 0.006–11.0s | 4.3–38.4s | 200 in 9.5–39.8s, or timeout at 30–60s |

▎ | Int custom domain (DNS-only, almost no traffic) | 0.04s–timeout | 12.0–16.7s | 200 in 18.4s, or timeout at 30s |

The int service gets almost no traffic and is equally affected, so this doesn't look like load on our side.

  1. Cloudflare reports errors from the origin. Production is proxied through Cloudflare, and clients get intermittent 520 Web server is returning an unknown error responses.

Ask

Please check the health of the Southeast Asia edge proxies handling these domains, especially TCP accept and TLS handshake latency during 04:00–06:00 UTC. Is there a known issue or a capacity problem on that edge? If it's isolated to particular edge nodes, can our domains be moved to healthy ones?

This affects a production API on the hot path of a live customer application, so we'd appreciate a quick response.

$10 Bounty

4 Replies

Railway
BOT

4 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 4 days ago


Anonymous
PRO

4 days ago

Same thing from India this morning, against the Singapore region, in bursts. It is not your app being slow.

Run these two lines while it feels slow:

curl -s -o /dev/null -w "app tls=%{time_appconnect}s total=%{time_total}s\n" 'https://YOUR_APP_DOMAIN/health'

curl -s -o /dev/null -w "google tls=%{time_appconnect}s total=%{time_total}s\n" 'https://www.google.com'

Then open your service HTTP logs for that same minute.

What we saw, on a deployment that was otherwise idle:

Google from the same laptop: about 0.2s

railway.com and our Singapore app, from that laptop: often 10–20s, most of it in the TLS handshake

The API container’s own log for the same request: about 50ms

A static file the server logged at 1ms still took many seconds on the client

So the handler had already finished. The missing seconds were between Bangalore and the container, and they came and went. Rough windows today (IST): about 10:00–10:30, 10:50–11:10, and 11:20–11:30. By 11:43 IST the same checks were back under 0.3s.

Nothing in the app or the database setting fixes that path. When your two curls match this split, send Railway those times plus the HTTP log line. That is the evidence for the latency on the edge.

Does that check help ?


Same thing from India this morning, against the Singapore region, in bursts. It is not your app being slow. Run these two lines while it feels slow: curl -s -o /dev/null -w "app tls=%{time_appconnect}s total=%{time_total}s\n" 'https://YOUR_APP_DOMAIN/health' curl -s -o /dev/null -w "google tls=%{time_appconnect}s total=%{time_total}s\n" 'https://www.google.com' Then open your service HTTP logs for that same minute. What we saw, on a deployment that was otherwise idle: Google from the same laptop: about 0.2s railway.com and our Singapore app, from that laptop: often 10–20s, most of it in the TLS handshake The API container’s own log for the same request: about 50ms A static file the server logged at 1ms still took many seconds on the client So the handler had already finished. The missing seconds were between Bangalore and the container, and they came and went. Rough windows today (IST): about 10:00–10:30, 10:50–11:10, and 11:20–11:30. By 11:43 IST the same checks were back under 0.3s. Nothing in the app or the database setting fixes that path. When your two curls match this split, send Railway those times plus the HTTP log line. That is the evidence for the latency on the edge. Does that check help ?

Anonymous
HOBBYOP

4 days ago

Thanks for confirming you also saw the same issue. I appreciate it 🙏.

So do we just wait an hope Railway fixes it? I see you're on the "Pro" plan, so I presume that means you're actually able to chat with Railway / get direct help from them to solve issues like this.

If it happens again today in an hour or so I guess I'll need to upgrade to "Pro" to ask them about it. First time seeing this kind of issue on Railway (been a user for over a year), so it hit me by surprise.

Thanks all.


Anonymous
HOBBYOP

3 days ago

Update: Oct 2 peak (03:45–06:20 UTC) was clean — ~0% 499s at the edge vs ~67% at the same time on Oct 1. No changes on our side, so it looks like it resolved on Railway's end. Would still appreciate confirmation of what happened.


Same thing from India this morning, against the Singapore region, in bursts. It is not your app being slow. Run these two lines while it feels slow: curl -s -o /dev/null -w "app tls=%{time_appconnect}s total=%{time_total}s\n" 'https://YOUR_APP_DOMAIN/health' curl -s -o /dev/null -w "google tls=%{time_appconnect}s total=%{time_total}s\n" 'https://www.google.com' Then open your service HTTP logs for that same minute. What we saw, on a deployment that was otherwise idle: Google from the same laptop: about 0.2s railway.com and our Singapore app, from that laptop: often 10–20s, most of it in the TLS handshake The API container’s own log for the same request: about 50ms A static file the server logged at 1ms still took many seconds on the client So the handler had already finished. The missing seconds were between Bangalore and the container, and they came and went. Rough windows today (IST): about 10:00–10:30, 10:50–11:10, and 11:20–11:30. By 11:43 IST the same checks were back under 0.3s. Nothing in the app or the database setting fixes that path. When your two curls match this split, send Railway those times plus the HTTP log line. That is the evidence for the latency on the edge. Does that check help ?

parthai-cto
PRO

3 days ago

Unfortunately I had my product demo during the same time. This was our first launch and we were seriously considering shifting away from railway because of this incident. I thought it had to do with my Hobby plan and I even upgraded to PRO as I had the demo. @railway we really need details and assurance of non occurrence of such incidents. On the status page also there is no update regarding this incident.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...