Lag on specific regions
kuhamaven
PROOP

3 months ago

Something is going on with the network connections from Railway. Noticed some users complaining, and after some tests, users from LATAM are getting response times of 22 - 30 seconds, while users outside of LATAM still get the usual 0.01 - 0.1 seconds.

What is going on? Such a huge lag is causing timeouts on all users.

Closed

78 Replies

kuhamaven
PROOP

3 months ago

Postman test from Argentina

image.png

Attachments


kuhamaven
PROOP

3 months ago

Requests on the server from online players outside of latam

image.png

Attachments


kuhamaven
PROOP

3 months ago

It's slowly normalizing


kuhamaven
PROOP

3 months ago

Still slow in LATAM, but at least numbers are going down


3 months ago

Thanks for the update 🙏


kuhamaven
PROOP

3 months ago

I'm getting 4-5 seconds on postman instead of 30


kuhamaven
PROOP

3 months ago

still not ms as it should, but its progress


kuhamaven
PROOP

3 months ago

Don't know what's going on, but I changed nothing on the server or anything


kuhamaven
PROOP

3 months ago

@Fragly Did something change or came up?


3 months ago

I'm figuring it out right now, sorry for the delay


kuhamaven
PROOP

3 months ago

users are being able to login again from LATAM


kuhamaven
PROOP

3 months ago

So huge progress W already


Anonymous
PRO

3 months ago

It seems to be back to normal for me


3 months ago

thanks for the updates


3 months ago

I've informed the team


kuhamaven
PROOP

3 months ago

got a bit laggy again


3 months ago

Kuha, where geographically are most of your users from? Or are they global?


kuhamaven
PROOP

3 months ago

I've gotten reports from:

Argentina

Chile

Mexico

Perú

Colombia


3 months ago

So South America generally?


kuhamaven
PROOP

3 months ago

Central America too


kuhamaven
PROOP

3 months ago

No reports from Europe or North America so far


3 months ago

Thank you 🙏 that helps a lot


systemrhc
HOBBY

3 months ago

Brazil kinda bad too


kuhamaven
PROOP

3 months ago

Anything I can get or relay, trust me I will!


kuhamaven
PROOP

3 months ago

And if you want me to run some tests, I can


3 months ago

I appreciate it, I'll let you know, thank you 🙂


3 months ago

Thank you for the info <:Pray:1360710432029147186>


deivipluss
HOBBY

3 months ago

The same thing happens to me, Peru!


3 months ago

It'd be helpful if someone experiencing slow requests could run an mtr or traceroute to 69.46.46.46


kuhamaven
PROOP

3 months ago

Spiked again to 36 seconds repsonse time


kuhamaven
PROOP

3 months ago

let me try


kuhamaven
PROOP

3 months ago

image.png

Attachments


3 months ago

Hmm could you install the mtr tool and try that instead?


devmedical
PRO

3 months ago

ping (15 pkts): min/avg/max/stddev = 69.4 / 92.9 / 164.2 / 28.9 ms, 0% packet loss (but notable jitter, σ≈29 ms).

traceroute reaches Miami cleanly and dies at CDN77 (ICMP filtered past it):

3 telefonicaglobalsolutions.com ~9 ms

7-9 *.mia03.atlas.cogentco.com (Miami) ~70-143 ms

10 vl221.mia-eq6-dist-2.cdn77.com ~70 ms

11-30 * * * (filtered, but 69.46.46.46 answers ICMP, ttl=54)

Key signal: public network from LATAM to your Miami edge is healthy (~93 ms, 0% loss), but HTTP TTFB to my service is 2–18 s (intermittent). My app's upstreamRqDuration is ~166 ms; the time is lost in the

edge→origin path. My service is deployed in us-west2 — looks like the edge(Miami)→us-west2 backhaul or the proxy is the bottleneck, not the LATAM→Miami public path. Service: web / project insightful-courage.

Slow request-IDs: 9l_uuPl7RyOvkhiKlt7tkg, 7aZv4ptZTZ2R2J4wWUN5dQ.


kuhamaven
PROOP

3 months ago

I'm on it


mativm02
PRO

3 months ago

this is from Argentina

image.png

Attachments


kuhamaven
PROOP

3 months ago

@Phineas

image.png

Attachments


kuhamaven
PROOP

3 months ago

Postman visible above as the latest request took 31 seconds


3 months ago

Thank you


3 months ago

We're investigating


kuhamaven
PROOP

3 months ago

image.png

Attachments


3 months ago

We have taken the Miami POP offline which was experiencing elevated packet loss.


3 months ago

It should be fixed now, and we will bring it back online once we the issue is resolved.


kuhamaven
PROOP

3 months ago

Okay... got 4.38 seconds now instead of 32


kuhamaven
PROOP

3 months ago

progress


mativm02
PRO

3 months ago

I think it got worse in our case :/


kuhamaven
PROOP

3 months ago

Yeah...


kuhamaven
PROOP

3 months ago

@Phineas spiking back to 7 and above


brunobollati
PRO

3 months ago

Im having the same issue. It's just too slow


devmedical
PRO

3 months ago

Update from Peru, isolating the remaining latency:

  • Client → edge atl1 (69.46.46.123): clean — ping 0% loss, avg 102 ms (RTT to Atlanta, expected).
  • App processing: healthy — gunicorn upstream time avg 86 ms (125 reqs, max 1.3 s), origin in us-west2.
  • But external TTFB is 2–3 s, spiking to 22 s.

So with Miami dropped, a clean client→atl1 path, and a fast origin, the ~2+ seconds unaccounted for are in the atl1→us-west2 edge↔origin backhaul — that's where the bottleneck moved. The Atlanta→us-west2 internal path looks congested/lossy. Happy to run more tests if useful.


3 months ago

The issue is with a backbone internet service provider, Arelion. We use them in US West so lots of internet traffic into US West is seeing issues. We're working on mitigating this.


brunobollati
PRO

3 months ago

Test made by my side

try1: ttfb=13.147113s total=13.147166s

try2: ttfb=7.557014s total=7.557089s

try3: ttfb=6.340120s total=6.340174s

try4: ttfb=15.468342s total=15.468453s

try5: ttfb=12.789899s total=12.789947s

try6: ttfb=1.077543s total=1.077596s


3 months ago

we've dropped Arelion AS1299 from our US-West region, you should start to see improvement shortly


mativm02
PRO

3 months ago

thanks, I'll keep you posted :)


kuhamaven
PROOP

3 months ago

Response times are three times longer now...


kuhamaven
PROOP

3 months ago

Wait, they are improving


mativm02
PRO

3 months ago

yeah, response times are a lot better right now


kuhamaven
PROOP

3 months ago

Yeap!


kuhamaven
PROOP

3 months ago

Users are able to log in finally


kuhamaven
PROOP

3 months ago

and response times are going back to normal


kuhamaven
PROOP

3 months ago

they are being just 0.1 seconds slower than usual


kuhamaven
PROOP

3 months ago

@Phineas Back again with this


kuhamaven
PROOP

3 months ago

24.78 seconds per request...


mativm02
PRO

3 months ago

same here


kuhamaven
PROOP

3 months ago

@Fragly


3 months ago

i let the team know


fvr1
PRO

3 months ago

same here (Chile)


3 months ago

Team is working on it now


kuhamaven
PROOP

3 months ago

Is the same thing from last time?


kuhamaven
PROOP

3 months ago

Pretty annoying having everything collapse every single week...


3 months ago

I have no idea, I believe the team is investigating the cause right now


3 months ago

will report back if I get any more info


pg
PRO

3 months ago

Are you opening an incident?


kuhamaven
PROOP

3 months ago

Well, was fixed and now it started lagging again...


3 months ago

Incident just called:


kuhamaven
PROOP

3 months ago

Cant it be redirected or changed like last time?


3 months ago

I haven't been told whether it even is the same issue as last time, although if it is then yes it could be redirected


ricardomacario
PRO

3 months ago

Our clients are in LATAM and they are all reporting the problem. We're in emergency mode.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...