Heavy packet loss between asia-southeast1 (Metal) and one Los Angeles host (AS23470 via NTT) since 23 Sep
rishipatel185
PROOP

10 days ago

Project: energetic-respect (fa509bfb-a5b3-4cab-b850-bb5d9cdb4208), environment production, service scriptaa-v2 (97f25a66-05ab-4bca-bfb0-e3909d58cd3d), region asia-southeast1-eqsg3a, Hobby plan, single replica, IPv6 egress off, current egress address 208.77.246.66.

What the service does: it polls two customer FTP servers that share one host, 104.194.8.126 (ReliableSite, AS23470, Los Angeles), on port 21 with passive data connections, and uploads finished documents back to the same host.

What changed and when (UTC):

  • Up to 23 Sep 2026 about 02:30, downloads from that host ran at 160 KB/s to 2.5 MB/s per file (measured from Railway network flow logs).
    • From 23 Sep about 15:00, downloads dropped to 20 to 45 KB/s, and a plain curl inside the container gets 9 KB/s from the same host. The same file downloads at 1.5 MB/s from an office connection in India.
    • From 23 Sep 19:00 to 24 Sep 00:00 the path was healthy again (about 780 KB/s, 81 MB moved) on the same deployment and the same egress address, then it degraded again and has stayed degraded.
    • The application code that talks to this host did not change across that boundary. Seven deployments since then (each with a new egress address) made no difference. The region has been asia-southeast1-eqsg3a for every deployment since at least 15 Sep.

Measurements taken on 25 Sep 2026 between 08:00 and 10:00 UTC:

  • From the ReliableSite Los Angeles looking glass (la1-lg.reliablesite.net, the same network as the remote host): ping to 208.77.246.66 answered 21 of 32 packets (34 percent loss, 222 ms); ping to 208.77.246.240, .65 and .67 answered 13 of 20; ping to a Vultr Singapore host (45.32.100.168) answered 16 of 16 at 190 ms. Traceroute from there to 208.77.246.66 goes ae-2.r26.lsanca07.us.bb.gin.ntt.net (129.250.3.90), ae-5.r29.osakjp02.jp.bb.gin.ntt.net (129.250.2.177), ce-3-3-2.a06.sngpsi07.sg.ce.gin.ntt.net (116.51.16.19, 234 ms), then no further replies.
    • From the Hurricane Electric Los Angeles looking glass (core3.lax1.he.net): ping to 208.77.246.66 answered 15 of 15 at 176 ms; ping to 208.77.246.240 answered 5 of 5; ping to 104.194.8.126 answered 5 of 5 at 4 ms.
    • Inside the container: a 100 MB download from lax-ca-us-ping.vultr.com ran at 3.5 MB/s and 0.8 MB/s on two tries; ping to the NTT Singapore edge 116.51.16.19 answered 20 of 20.
    • Inside the container, TCP counters since boot: TCPOFOQueue 36,850 of 415,824 inbound segments, TCPDSACKOldSent 2,186, TCPTimeouts 214, TCPSynRetrans 161 of 994 connection attempts. The control socket to 104.194.8.126:21 sat at cwnd 2 with a 2.08 s retransmit timer; a data socket had a 4.41 s retransmit timer. Receive side is clean: TCPOFODrop 0, TCPRcvQDrop 0, TCPBacklogDrop 0, no cgroup CPU throttling, 212 MB of 8 GB memory in use.
    • Railway network flow logs for ingress from 104.194.8.126 show l4LatencyMs of 1,090 to 3,311 ms during transfers and dropCause values TCP_INVALID_SEQUENCE, TCP_ACK_UNSENT_DATA and TCP_AOFAILURE.
    • Uploads to the same host are affected too: a 9 KB file now takes 31 to 40 s end to end (it took 11 to 13 s before 23 Sep), though every upload verifies correctly.

What this looks like from our side: packets between 208.77.246.0/24 in asia-southeast1 and AS23470 are being lost or delayed in both directions on the path that runs through NTT (AS2914), while the same Railway addresses are reachable cleanly from Hurricane Electric in Los Angeles and the same remote host is reachable cleanly from elsewhere.

Questions:

  1. Is there a known issue or a recent change on the asia-southeast1 ingress or egress path towards NTT since 23 Sep 2026 15:00 UTC?
  2. Can you check for loss on the NTT peering or transit for 208.77.246.0/24, or on the host that holds 208.77.246.66?
  3. If this is outside Railway's control, is there a way to steer this service's traffic to a different upstream, or is a region move the only option?

Happy to run any test you need from inside the container. Thank you.

Awaiting Conductor Response$10 Bounty

2 Replies

Railway
BOT

10 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 10 days ago


zensteagarden
PRO

7 days ago

Your measurements already isolate this. It is not the FTP client, the container, or the specific egress address.

  1. Known Railway change on asia-southeast1 since 23 Sep 15:00 UTC?

No. Railway status for that day is the EU West build queue only. Nothing posted for asia-southeast1-eqsg3a ingress/egress in this window.

What did change on that path: NTT GIN lost Asia capacity after multiple subsea cuts (from ~21 Sep 02:20 UTC), with Singapore / HK / Taiwan transit hit hardest, plus Typhoon Dujuan on 21 Sep. Quad9 filed that at 23 Sep 11:04 UTC and still had it degraded. Cloudflare incident 8wmvkv5jkf15 is the same congestion between Tokyo and Singapore; last public update 25 Sep 08:50 UTC, still mitigating.

Your 23 Sep 19:00–24 Sep 00:00 healthy window on the same deployment and same .66 is that congestion oscillating. Seven new egress IPs in 208.77.246.0/24 would not escape it.

  1. Loss on NTT toward 208.77.246.0/24 / host .66?

I cannot read Railway peering counters. Public routing plus your probes are enough:

208.77.246.0/23 is Railway AS400940 Singapore space. Listed upstreams include NTT AS2914 and PCCW AS3491.

ReliableSite LA (same network as 104.194.8.126) loses packets to .66 and traceroute dies after NTT Singapore 116.51.16.19.

Hurricane Electric LA reaches both .66 and 104.194.8.126 cleanly (~4 ms to the FTP host).

Inside the box: cwnd 2, 2–4 s RTOs, TCPOFOQueue ~9% of inbound segments, flow-log l4LatencyMs 1.1–3.3 s with TCP_INVALID_SEQUENCE / TCP_ACK_UNSENT_DATA / TCP_AOFAILURE. Receive drops are zero, no CPU throttle. That is middle-path loss/reordering, not a dead host.

  1. Steer off NTT, or only a region move?

Hobby cannot pick an upstream. Static outbound IPs are Pro. IPv6 egress will not help an IPv4 FTP peer on :21. Redeploying in asia-southeast1-eqsg3a keeps you on 208.77.246.0/23 and the same NTT-toward-AS23470 path.

Customer fix: move scriptaa-v2 to us-west2 (California). The peer is ReliableSite Los Angeles. Region change on Metal does not require a domain change. Optional control: add one extra replica in us-west2 first, download the same file and time the 9 KB upload. If you are back in the MB/s class, the path diagnosis is confirmed.

If Railway can announce 208.77.246.0/24 away from NTT toward AS23470, that would also fix Singapore-region copies. Until that happens, us-west2 is the lever you actually have.

Happy to run a specific in-container check if useful (mtr -w -c 50 104.194.8.126, ss -ti on the data socket, or a timed curl to a nearby LA host vs the FTP host).


rishipatel185
PROOP

7 days ago

Thanks!!


Welcome!

Sign in to your Railway account to join the conversation.

Loading...