Edge sin1: TLS handshake 10–28s / timeouts to *.up.railway.app — app responds in ~40ms
arkomr
PROOP

15 hours ago

Since ~04:20 UTC on 2026-10-01 our users in Thailand see very slow page loads and "connection timed out" errors. Status page shows all green.

The app itself is healthy: railway logs --http shows every request served in 30–130 ms (including DB). The delay is before the request reaches our container — in the TLS handshake to the Railway edge (x-railway-edge: sin1). It affects a static frontend service too, so it is not app/DB related.

curl measurements from Bangkok (time_appconnect = TLS done):

| UTC | URL | TLS | total | x-railway-request-id |

| --- | --- | --- | --- | --- |

| 04:43 | rentalapi-production-d3f4.up.railway.app/health | 28.06s | 29.72s | — |

| 04:43 | rentalfrontend-production.up.railway.app/ | 23.65s | 28.61s | — |

| 04:43 | krp2-api.koder3.com/health (via Cloudflare) | — | 3.03s (up to 54.8s earlier) | TKxVs353QMShTM7TDcO5xA |

| 04:48 | rentalapi-production-d3f4.up.railway.app/health | timeout | 60s (no response) | — |

| 04:49 | rentalapi-production-d3f4.up.railway.app/health | 10.29s | 15.04s | q4-pJgocSvSNidpfAXC71g |

| 04:49 | rentalapi-production-d3f4.up.railway.app/health | 13.73s | 19.28s | yRSvqdXBRviqY2ONDcO5xA |

| 04:50 | rentalfrontend-production.up.railway.app/ | 14.85s | 17.74s | 3wqrtXUzSoersRTlJH0Vcg |

Same request in our http logs, e.g. GET /health 200 47ms. WebSocket /api/v1/ws also hangs ~25s before dropping.

DNS and TCP connect are fast (<0.1s); the TLS handshake at the edge is the slow part. Intermittent: some requests complete in ~1s.

Project ID: 0aad612f-b9f2-444c-b11b-77ad0595b10b (environment: production, region: Southeast Asia). Services: api, rental_frontend.

Similar symptoms on 2026-09-30 from ~07:40 UTC (public networking, SEA).

Could you check the sin1 edge? Thanks.

$20 Bounty

6 Replies

Railway
BOT

15 hours ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 15 hours ago


willscottrod
PRO

15 hours ago

Seeing the same issue as well


plumsydev
PRO

15 hours ago

Hey, Railway servers are experiencing a huge lag issues lately.

Please remain calm and wait for the team to answer questions on that.


markreo
HOBBY

14 hours ago

A few minutes ago, from around 04:50 to 05:00 UTC, after my customer reported the issue, all of our APIs became significantly slower.

The Railway dashboard (railway.com) was also taking a very long time to load and occasionally returned errors. However, the Railway status page showed all services as operational (green).


arkomr
PROOP

14 hours ago

Update: it came back. Briefly normal ~05:05–05:14 UTC (TLS ~0.1s), then degraded again from ~05:18 UTC: TLS 6–57s, one 60s timeout, one 502 from the edge (our http logs show no error; /health served in 35–62ms, p95 156ms across 38 requests). Logs also show 499s — edge dropping the connection before our app could respond. Please share a root cause — this is the second day in a row (2026-09-30 too).


bishhmalaysia
HOBBY

11 hours ago

Same here, seen from a different client network. Our API is a custom domain on Railway (Singapore edge, 69.46.46.105). It is called by our website's server functions on Vercel in Singapore (sin1, AWS ap-southeast-1). On 1 Oct 2026:

  • New TCP/TLS connections timed out after 10 s (undici UND_ERR_CONNECT_TIMEOUT). Some showed connect ETIMEDOUT 69.46.46.105:443 or other side closed with 0 bytes read.
  • About 372 timeouts that day, against 4 in the previous 30 days.
  • Per 10 minutes (UTC): 04:20 = 23, 04:30 = 103, 04:40 = 31, 04:50 = 33, 05:00 = 0, 05:10 = 10, 05:20 = 70, 05:30 = 95, 05:40 = 0, 05:50 = 16. Last one at 05:58:46. We saw the same pause around 05:05 to 05:15 that you did.
  • At the worst points, roughly 40% of connection attempts failed.
  • 17 HTTP 502 "upstream error" responses reached us (01:14, 01:17, then between 04:36 and 05:31 UTC). None of them appear in our service's HTTP logs.
  • Our service's HTTP logs have no rows at all from 04:14:03 to 04:15:02 UTC, while the app was serving requests.
  • The service itself was healthy throughout: zero 5xx, and requests that did arrive were answered in 4 to 74 ms.
  • The same domain worked fine from a Malaysian ISP (via the hkg1 edge) and from Vercel's US build machines.

So for us it lasted about an hour and a half, not a few minutes. We also hit a separate event on 30 Sep around 07:35 to 07:56 UTC.

Railway team: could you share the root cause, why the status page stayed green, and what is being done so this does not repeat? Thank you.


matchup-tech
HOBBY

6 hours ago

Same issue here, also on the sin1 edge, and it was still happening after the window reported above.

Our production API is a custom domain on Railway (Southeast Asia, x-railway-edge: sin1, Server: railway-hikari). Our mobile app's users and our own machines in Malaysia see the TLS handshake to the edge stall, while the app answers in milliseconds:

  • 2026-10-01 07:11:32 UTC, curl from Malaysia: TCP connect 0.037 s, TLS handshake 14.8 s, total 20.8 s. Our Railway HTTP log shows that same request at 12 ms total / 12 ms upstream. Nine more requests in the same run took 0.08 to 3.1 s.
  • 05:00 to 05:30 UTC: repeated stalls from a second Malaysian network. Our app's 30 s client timeout fired twice on cold launch, and one request was logged as a 499 at 05:20:28 UTC.
  • No deploys, restarts or errors on our side at the time. 2 replicas, sleep disabled. DNS and TCP connect are fast; only the TLS handshake at the edge is slow.

Could the Railway team share the root cause and whether sin1 is now stable? Thank you.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...