20 days ago
Service: ABC-Teams-Backend (project ABCcount-Teams, production environment) Affected domain: api.abccount-teams.com Working domain on the same service: abc-teams-backend-production.up.railway.app
Summary
Requests to our custom domain api.abccount-teams.com fail during the TLS handshake from a Comcast residential connection in Houston, Texas. Requests to the same service over its *.up.railway.app domain, from the same machine on the same connection seconds apart, succeed normally.
The same custom domain works from a mobile hotspot on a different carrier. So the certificate appears to be present on some edge nodes and absent or invalid on others.
This has been reproducible for over a week. A previous ticket on this issue was closed; the problem has not gone away. We have since removed and re-added the custom domain, which issued a new CNAME target, and the behavior is unchanged.
Reproduction
All commands run from the same Windows machine, same network, within the same minute. curl 8.9.0, Schannel TLS backend.
Fails — custom domain
$ curl -v https://api.abccount-teams.com/health
- Host api.abccount-teams.com:443 was resolved.
- IPv4: 69.46.46.127
- Trying 69.46.46.127:443...
- Connected to api.abccount-teams.com (69.46.46.127) port 443
- schannel: disabled automatic use of client certificate
- schannel: next InitializeSecurityContext failed: SEC_E_INVALID_TOKEN (0x80090308)
- The token supplied to the function is invalid
- closing connection #0
curl: (35) schannel: next InitializeSecurityContext failed: SEC_E_INVALID_TOKEN
The connection opens on port 443 and then fails at the first handshake step. SEC_E_INVALID_TOKEN means the server's reply was not a valid TLS message.
Succeeds — same service, railway.app domain
$ curl -v https://abc-teams-backend-production.up.railway.app/health
- Connected to abc-teams-backend-production.up.railway.app (69.46.46.17) port 443
- schannel: disabled automatic use of client certificate
GET /health HTTP/1.1
< HTTP/1.1 200 OK
< Server: railway-hikari
< x-railway-request-id: EN32f27pTKScl6vOvEOTxA
< x-hikari-trace: iah1.k00j
< x-railway-edge: iah1
{"ok":true}
Succeeds — custom domain, different network
From a mobile hotspot, same machine, same binary:
$ curl -v https://api.abccount-teams.com/health
< HTTP/1.1 200 OK
< Server: railway-hikari
< x-railway-request-id: rEBnm8c0QiGw5MYwN8N_Fg
< x-hikari-trace: iah1.2nva
< x-railway-edge: iah1
{"ok":true}
Note that the working hotspot request and the failing home request both report edge iah1.
What we have ruled out
DNS. Both our ISP's resolver and Google's return the same record, and it matches what the Railway dashboard asks for.
$ nslookup api.abccount-teams.com 8.8.8.8
Name: mgwwikpc.up.railway.app
Address: 69.46.46.127
Aliases: api.abccount-teams.com
Routing. tracert reaches the host in 11 hops with no loss. The two timed-out hops in the middle are routers not answering ICMP.
TLS version. Forcing TLS 1.2 (--tlsv1.2 --tls-max 1.2) produces the same failure.
Certificate revocation checking. --ssl-no-revoke produces the same failure.
Local proxy or interception. --proxy "" produces the same failure. Other HTTPS hosts, including railway.app itself, work normally from this machine.
SNI as the only variable. The failing and succeeding requests go to Railway edge addresses in the same /24, from the same client, seconds apart. The difference is the hostname sent in SNI.
Stale DNS or configuration on our side. We removed and re-added the custom domain. Railway issued a new CNAME target (z3bvuymg → mgwwikpc) and a new _railway-verify TXT record. Both were updated in Netlify DNS and verify correctly from a public resolver. The dashboard shows the domain as issued. The handshake failure is unchanged.
Impact
api.abccount-teams.com is the API for our entire B2B product. Our Chrome extension and Windows desktop app both talk to it exclusively. For any user on an affected network, the product does not work at all — and because the failure happens during the handshake, the clients cannot distinguish it from an authentication problem.
We are about to onboard a hundred pilot companies, mostly law firms and medical practices. If one ordinary residential connection is affected, some proportion of corporate networks will be too, and we have no way to diagnose those remotely.
What we are asking
Whether the certificate for api.abccount-teams.com is present and valid on every edge node serving 69.46.46.0/24, or only some.
Whether anything in the previous domain removal and re-add left inconsistent state across edges.
Whether there is a way for us to detect this condition ourselves, so we can tell customers "our provider has an edge problem" rather than guessing.
Request IDs
Result Request ID Edge
Success, railway.app domain, home network EN32f27pTKScl6vOvEOTxA iah1
Success, custom domain, hotspot rEBnm8c0QiGw5MYwN8N_Fg iah1
Fail, custom domain, home network no ID — handshake never completed
2 Replies
20 days ago
The domain's certificate is valid, DNS is correctly propagated to the expected CNAME target, and traffic routes are installed on our edge, so from our side everything is healthy. The pattern you're describing, where TLS fails from one ISP but succeeds from another to the same edge node, matches middlebox or ISP-level filtering on the network path between the client and our edge. Two tests can narrow it down: use curl --resolve to force the hostname onto a different edge IP (success means IP-based filtering), and use openssl s_client -servername to send the custom domain's SNI to the same IP the generated domain resolves to (a hang on the custom SNI but not the generated one means SNI-based filtering). We cannot assign or exclude specific edge IPs, and redeploys or region changes do not change them, so the lasting fix for affected networks is placing a CDN or reverse proxy (such as Cloudflare) in front of the custom domain so that clients reach an IP range their ISP does not filter.
Status changed to Awaiting User Response Railway • 20 days ago
Railway
The domain's certificate is valid, DNS is correctly propagated to the expected CNAME target, and traffic routes are installed on our edge, so from our side everything is healthy. The pattern you're describing, where TLS fails from one ISP but succeeds from another to the same edge node, matches middlebox or ISP-level filtering on the network path between the client and our edge. Two tests can narrow it down: use `curl --resolve` to force the hostname onto a different edge IP (success means IP-based filtering), and use `openssl s_client -servername` to send the custom domain's SNI to the same IP the generated domain resolves to (a hang on the custom SNI but not the generated one means SNI-based filtering). We cannot assign or exclude specific edge IPs, and redeploys or region changes do not change them, so the lasting fix for affected networks is placing a CDN or reverse proxy (such as Cloudflare) in front of the custom domain so that clients reach an IP range their ISP does not filter.
20 days ago
Ran both tests. Test 1: forcing the custom hostname onto 69.46.46.17 (the IP the generated domain uses successfully) still fails the handshake, so it isn't IP-based filtering. Test 2 is conclusive — same IP, same second, only the SNI differs:
-servername api.abccount-teams.com → packet length too long, no peer certificate, handshake dies after 5 bytes read.
-servername abc-teams-backend-production.up.railway.app → full chain, CN=*.up.railway.app, valid.
That's SNI-based filtering on the path between us and your edge. Your diagnosis was right and it's not a Railway problem. We'll be putting Cloudflare in front of the custom domain as you suggested. Thanks for the specific tests
Status changed to Awaiting Railway Response Railway • 20 days ago
Status changed to Solved Railway • 20 days ago