Custom domain repeatedly unreachable — Railway edge target dying
journeylink-sundog
HOBBYOP

2 months ago

  • Our custom domain stopped accepting connections

    entirely — TCP handshakes never completed (curl: connect=0.000000s, hangs to

    timeout). Users saw "Can't reach the server."

  • Our container was completely healthy throughout: scheduled background jobs

    kept logging on time for the full outage window, and the DB stayed reachable.

    No crash, no OOM, no restart.

  • The service's generated domain served perfectly the entire time — GET /api/version returned in ~200ms with a clean TLS handshake (301 via our canonical-host redirect).

  • The custom domain's CNAME target that Railway had assigned, resolved to a Railway-owned edge IP that refused/hung connections. whois confirms

    these are Railway IPs (RLWY-HIKARI-01, Railway RC-1550). So the live service

    moved to a different Railway edge IP (69.46.46.82) and the edge behind the

    custom domain's target was dead.

  • A Railway service RESTART did NOT fix it (we polled for 2+ minutes after) —

    confirming the fault is in the edge/routing layer, independent of the container.

Solution/fix? Our current solution: Delete and re-add the custom domain and wait for it to come live, not very sustainable. Thanks!

$10 Bounty

2 Replies

Railway
BOT

2 months ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 2 months ago


brezzy1337
PRO

2 months ago

Yeah that issue sounds very frustrating, edge network problems can be a real pain, and are hard to reproduce externally.

There is one thing you can try that will scope in the solution and that is bypass DNS and force the connection to the edge IP you've already confirmed is live:

dig +short yourservice.up.railway.app

curl -sv -- resolve yourdomain.com:443 :< that-ip>https://yourdomain.com/api/version

If that succeeds while the domain itself still hangs, your container and Railway's edge are both fine - the problem is which IP your DNS hands out.

The usual cause is a DNS provider publishing a flattened or proxied record instead of a Live CNAME: it resolves Railway's CNAME target once, caches the result as its own

A (or AAAA) record, and keeps serving that IP after Railway rotates edges. That fits everything you saw - generated domain fine (it resolves live), container fine, restart useless, and delete-and-re-add working because it forces a re-flatten.

To separate the A and AAAA cases - note these run without -- resolve, since the point is to see what DNS actually returns per family:

curl -sv -4 https://yourdomain.com/api/version

curl -sv -6 https://yourdomain.com/api/version

If v4 works and v6 hangs, you have a stale AAAA.

Hope this narrow's down the issue, it very interesting issue so thanks for posting and looking forward to an update, best regards and happy building 😎


vaj31
FREE

2 months ago

Force an Edge Sync Without Deleting the Domain (Fastest Fix)If this happens again and you need to restore connectivity immediately without blowing away your domain settings or risking SSL rate limits:Scale Instances Up and Down:Go to Service Settings $\rightarrow$ Scaling $\rightarrow$ temporarily change Replicas from 1 to 2. Once the second instance is active, scale it back down to 1. This triggers a full upstream discovery update across Railway’s internal Envoy proxy mesh, forcing the edge to rebuild its routing tables. Toggle Environment Variables:Adding or changing a non-essential service environment variable (e.g., FORCE_EDGE_SYNC=1) triggers a full control-plane deployment cycle rather than a simple container restart


Welcome!

Sign in to your Railway account to join the conversation.

Loading...