Edge dropped SNI bindings for all custom domains during scheduled cert renewal (3rd domain incident in 5 days)
lilsquizy19
PROOP

a month ago

Our production custom domains (generatorpesen.ru, admin.generatorpesen.ru) went down

on Jul 15 at ~04:53-05:09 UTC. TLS handshakes were RESET (RST before ServerHello,

"no peer certificate available") on all edge anchors, port 443 and 80.

Root cause correlation: CT logs (certspotter) show your scheduled certificate renewal

issued new certs exactly at that time:

  • 04:53:11 UTC generatorpesen.ru
  • 04:58:42 UTC admin.generatorpesen.ru
  • 05:09:36 UTC generatorpesen.ru (again)

Previous certs were from Jun 10, so this was a scheduled renewal. The certs were

ISSUED but the edge network dropped the SNI bindings for both domains afterwards.

Evidence at the time of the incident:

  • openssl s_client with our SNI -> connection reset on BOTH anchors

    (69.46.46.67 / lyhi4su0 and 69.46.46.92 / i6wtcdmg)

  • The anchors themselves were healthy: wildcard *.up.railway.app cert served fine

  • Dashboard showed both domains green/active, DNS records (CNAME/ALIAS + both

    _railway-verify TXT) intact and correct

  • Postgres TCP proxy in the same project kept working, app was healthy

    (generated .up.railway.app domain worked immediately)

What we did:

  • Toggling the target port did NOT force the edge to re-register the SNI

  • Deleting and re-adding the domains + updating DNS to the newly assigned targets

    fixed admin.generatorpesen.ru within ~3 minutes (now on ci27s99b)

  • The root domain generatorpesen.ru (re-added, now assigned evggm12x, DNS updated,

    dashboard checkmarks green) is still NOT getting its certificate deployed after

    20+ minutes. We suspect the Let's Encrypt failed-validation rate limit

    (5/hostname/hour) was exhausted by the failed renewal attempts this morning.

This is our THIRD custom-domain incident in 5 days:

  • Jul 11: admin cert renewal silently failed (missing _railway-verify TXT on a

    grandfathered domain), cert expired

  • Jul 12: anchor i6wtcdmg lost the SNI record for admin during nightly edge

    rebalancing while the neighbor anchor served it fine

  • Jul 15: this incident, both domains, all anchors

Requests:

  1. Please investigate the renewal -> edge deploy pipeline that drops SNI bindings.

  2. Please check/unblock certificate issuance for generatorpesen.ru

    (binding target evggm12x.up.railway.app) - production is still down.

  3. Advise whether we need to change anything on our side to survive the next

    scheduled renewal.

$20 Bounty

1 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway about 1 month ago


I’m able to access your sites just fine. It may be due to an ISP or regional block. Try using. VPN to access your site, or proxy your domains with Cloudflare’s DNS.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...