Incident 17 Feb ~01:15 UTC

6 months ago

I'm not seeing any discussions related to this in the Discord..

Around 1-2 hours ago, I was notified of 2 things:

  • Hostname/IP does not match certificate's altnames: Host: unproxied-custom-domain-here. is not in the cert's altnames: DNS:*.up.railway.app

  • HTTP 520 responses for Cloudflare proxied custom domains

Looking further, there was also a spike in latency

This affected all of my services, across multiple projects, and all of the deployments were healthy w/o redeploys

The status checker I'm using here is hosted in AMS

Incident started around 01:15 UTC and ended around 01:25 UTC

Screenshot_20260217-045625.png

Screenshot_20260217-045701.png

Screenshot_20260217-045719.png

Solved

40 Replies

6 months ago

What do your http logs say?


6 months ago

Were all green, let me check again


6 months ago

Scrolling through the HTTP logs on a service that got the certificate issue, it's all 200s


6 months ago

Link


6 months ago

noticeable gap (utc+2 tzs)

Screenshot_20260217-050507.png

Attachments


6 months ago

linked, should be this one


6 months ago

To clearly clarify, this is no longer an issue


6 months ago

Screenshot_20260217-050657.png

Attachments


6 months ago

I'm not seeing anything on our end here.


6 months ago

in case missed


6 months ago

Yes, I've looked at our monitoring for the time you gave.


6 months ago

another example, cloudflare proxied

Screenshot_20260217-051204.png

Screenshot_20260217-052048.png


6 months ago

Not seeing anything on our end for that either.


6 months ago

I mean, could it be an observability issue? I don't think Cloudflare lies about 520s


6 months ago

I am honestly not sure.


6 months ago

Not a new issue either I believe

Screenshot_20260217-053634.png

Attachments


6 months ago

I'm also not seeing other users report this 🙁


6 months ago

"maybe I'm just not like other users"


6 months ago

think about the times where I called out an outage first (at least in the discord)!!


6 months ago

Haha I'm not sure what you want me to say to that.


6 months ago

didn't expect you to


6 months ago

I'm fine with leaving it at this, no SLA or critical services on my end


6 months ago

may be worth looking into (deeply) though..


6 months ago

I've looked, I don't see anything from any of our monitors.


6 months ago

what about now!


haayhappen
PRO

6 months ago

I'm also observing this (second time today a couple minutes ago) - first time was 8:20 UTC


mai1015
HOBBY

6 months ago

image.png

Attachments


mai1015
HOBBY

6 months ago

wow so serious the server down


6 months ago

^


futsy
HOBBY

6 months ago

Also seeing this as well


duxsec
HOBBY

6 months ago

Interested to see what this caused.


futsy
HOBBY

6 months ago

I've been seeing similar behaviour just now.

UTC timeline of the recent flapping I observed (triggered by my monitoring system; which is also monitoring services outside of Railway as well and only Railway was impacted:

08:21:35 — DOWN

08:22:56 — UP

08:24:58 — DOWN

08:26:18 — UP

09:40:46 — DOWN

09:44:44 — UP

09:58:27 — DOWN

09:59:47 — UP

10:01:53 — DOWN

During this window I did some diagnostics:

DNS resolution was normal and stable (single IPv4 A-record; no IPv6).

TCP connect to port 443 succeeded immediately.

The failure happened during TLS negotiation: curl sent the TLS ClientHello, then timed out waiting for the server’s ServerHello (connection established, but the TLS handshake response didn’t come back).

Shortly after, the alert cleared by itself. I re-ran the same diagnostics and TLS completed normally (TLS 1.3), ALPN negotiated HTTP/2, and the health request returned 200.

Note I am NOT behind CF proxy.


pepijn
PRO

6 months ago

Same here


6 months ago

"Same" and "this" meaning you all have seen a cert for a service domain instead of a custom domain?


Railway
BOT

6 months ago

Hello!

We've escalated your issue to our engineering team.

We aim to provide an update within 1 business day.

Please reply to this thread if you have any questions!

Status changed to Awaiting User Response Railway 6 months ago


6 months ago

Happy? 😆

image.png

Attachments


6 months ago

I win


6 months ago

<:proud:800594749714071553>


6 months ago

Hello,

We have provisioned more capacity in regards to our edge network in the EU-West region. You should no longer see these blips of incorrect certificates.


haayhappen
PRO

6 months ago

Thank you.. how can this be caught automatically next time?


Status changed to Awaiting Railway Response Railway 6 months ago


6 months ago

We are going to alert on the graph I shared above, and long term, we are well underway on a complete rewrite of the network control plane that will outright solve for this, and solve for many other potential issues.


Status changed to Awaiting User Response Railway 6 months ago


Status changed to Solved brody 6 months ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...