3 months ago
monitors went down showing "socket hang up" just now
91 Replies
3 months ago
might've been resolved on its own?
3 months ago
back green but still...
3 months ago
connection err closed something something when I tried to fetch my service's healthcheck endpoint
3 months ago
3 services were/are effected so far, no new deployments
3 months ago
will look around
3 months ago
I just had it happen on railway.com as well
3 months ago
I'm in eu ams
affected services whose healthchecks failed externally
3 months ago
Seeing latency spikes
3 months ago
both screenshots are of services linked above
Attachments
3 months ago
not seeing the socket hang up anymore but still seeing slight latency spikes
3 months ago
socket hang up again
3 months ago
may I ask which ip/isp you're getting these from? can dm too
3 months ago
502 bad gateways, connection reset by peer
3 months ago
69.46.46.14:443: i/o timeout
3 months ago
constantly, like right now?
3 months ago
upstream error
Attachments
3 months ago
uhh not super constantly but quite consistent
3 months ago
is there any response body?
3 months ago
upstream error
3 months ago
everything seems to be stable the last 30m, but might just be luck
3 months ago
let me know if you see one in the next 5min or so
3 months ago
instability is back...
3 months ago
unexpected EOF
3 months ago
read tcp 172.18.0.8:33922->66.33.22.216:443: read: connection reset by peer
3 months ago
ok, I know why you see this now
3 months ago
wait what
3 months ago
Oh I see, the origin/railway proxies (on the 66.33.22.0/24 prefix) were just deployed
3 months ago
How aggressively are you monitoring those? Do you keep the connection open?
3 months ago
I don't think so, minutely checks
3 months ago
I'm trying to access the failing URLs from my browser, I get hit with a conn refused (or alike, browser autorefreshes) and afterwards it seems to work fine
3 months ago
could you send me the failing URL in here/over DMs? (does it resolve to an address within 66.33.22.0/24)?
3 months ago
I'll send you the ones my monitors are annoying me about, do want to note that it's intermittent so far
3 months ago
seeing similar intermittent down behavior
3 months ago
happened again around 40 minutes ago
3 months ago
seeing timeouts, question mark
3 months ago
do you know if its the 66.33.22 IPs or 69.46.46 IPs you're seeing timeouts on?
3 months ago
69.46.46.58
3 months ago
resolved ""itself"" as of now but
Attachments
3 months ago
ok, that would do it. bgp should reconverge
Attachments
3 months ago
Let me know if you see anything else.
3 months ago
Was this from a monitor out of interest? Is the monitor hitting the "ams1" POP? https:///.railway/cdn-trace
Yeah, it was a monitor hitting my own API. The entire app was running into a Cloudflare timeout error, but the Railway metrics were still showing everything as healthy and operating normally so idk what was that
3 months ago
https://discord.com/channels/713503345364697088/1511663784127762483/1511663784127762483 seeing some latency spikes
3 months ago
69.46.46.58
Attachments
3 months ago
Could you run a traceroute/mtr to that IP from the monitor's network/server if possible? 🙏
3 months ago
not consistent latency... mtr just shows ams eq6 right now, but ill keep retrying
3 months ago
some cloudflare proxied requests are timing out completely
3 months ago
I'd like to see the hops you take before eq6
3 months ago
2.|-- sre02.gs.core.blackgate.nl 0.0% 10 4.6 4.7 4.3 6.0 0.5
3.|-- 100.65.1.14 0.0% 10 4.3 5.0 3.8 11.2 2.2
4.|-- 100.65.0.161 0.0% 10 8.6 5.6 3.9 9.8 2.0
5.|-- 81.20.68.161 0.0% 10 4.6 5.8 4.5 10.0 1.8
6.|-- ae-7.r23.amstnl07.nl.bb.gin.ntt.net 0.0% 10 84.0 85.5 83.4 90.8 2.5
7.|-- ae-18.a01.amstnl07.nl.bb.gin.ntt.net 0.0% 10 8.6 7.4 4.7 14.1 3.1
8.|-- 81.20.68.138 0.0% 10 3.4 4.3 3.3 10.5 2.2
9.|-- vl221.ams-eq6-dist-2.cdn77.com 0.0% 10 4.5 5.1 4.5 6.5 0.6
10.|-- 69.46.46.70 0.0% 10 4.9 5.0 4.3 7.6 0.93 months ago
https://utilities-us-east.up.railway.app/raw
What is the x-railway-edge header you see here?
3 months ago
railway/europe-west4-drams3a
3 months ago
What about here? https://utilities-us-east-cf-proxied.railway.com/raw
3 months ago
same edge val, cf-ray ending in -ams
3 months ago
Ty
3 months ago
timing out for me
3 months ago
Attachments
3 months ago
and... not anymore
3 months ago
No timeouts on https://utilities-us-east.up.railway.app/raw right - just via Cloudflare?
3 months ago
timeouts only observed via cloudflare so far
3 months ago
latency spikes on non-cloudflare still a thing but my monitors havent tripped on them
3 months ago
cf-ray became lhr, then timed out the following refresh
(and once again I am able to fetch it fine..)
3 months ago
I'm assuming you disabled the ams pop
Attachments
3 months ago
Nope it's still up, do you see that on the non-CF domain as well?
3 months ago
I do not
Attachments
3 months ago
I have a suspicion on what it could be. I'm going to disable something and we can see if that resolves it.
swag42dev
 railway/europe-west4-drams3a
3 months ago
Is your domain also proxied by Cloudflare?
phin
Is your domain also proxied by Cloudflare?
3 months ago
no
3 months ago
I've disabled something in AMS - let me know if you see improvement over the next half hour or so
phin
Is your domain also proxied by Cloudflare?
3 months ago
Domain management is on CF, but this particular domain is not proxied.
3 months ago
still seeing these (or should I just wait a bit more)
3 months ago
What latency is that tracking? ICMP or HTTP?
3 months ago
HTTP
phin
May I know the domain or service this is happening to?
3 months ago
project/b1f6fd55-ba7c-423a-8784-21ae43029bab/service/88da5b4b-99c2-4e7d-8946-b51c75496f9f
3 months ago
still going cf ams -> hikari lhr
3 months ago
non-proxied latency spikes still occurring
3 months ago
Is there a specific domain this is occuring on or is it all of your domains?
3 months ago
I'll dm
3 months ago
This is starting to affect other deployments.
The problem is becoming widespread.
Attachments
3 months ago
Hey.
Same here.
Using a VPN, we tested a few other edges.
In a nutshell, railway/europe-west4-drams3a is definitely troubles for us. us-west2 works fine as far as we can tell.
3 months ago
Monitors tripped over *.up.railway.app timeout
3 months ago
so it's no longer necessarily CF specific
3 months ago
I'm also seeing latency spikes from Singapore
(Singapore Railway service -> EU-W Railway service over pubnet)
3 months ago
I'm guessing something was changed about 18 minutes ago? Latency seems better
swag42dev
 This is starting to affect other deployments. The problem is becoming widespread.
3 months ago
I’m having the same issue on my Railway-hosted Node/NestJS backend, but Railway support has not responded yet
3 months ago
There's still few latency spikes going on, but not often
3 months ago
Interesting. All of those domains have pointed to the old edge/66.33.22.0/24 for about 8 hours.