a month ago
Production impact: This is a live production application (recently launched, actively serving users). Every database query from our production app server pays ~180ms of avoidable network latency, making uncached API responses take 4-8 seconds. We would appreciate prioritized handling.
Our Postgres service (Postgres-tNcb, europe-west4 Amsterdam, TCP proxy sakura.proxy.rlwy.net:25896) shows consistently high latency from our app server at Hetzner Nuremberg, Germany.
The problem:
- Client: Hetzner Nuremberg (IP 128.140.97.112, AS24940)
- RTT to anycast edge (66.33.22.0/24): consistently ~180ms
- Traceroute shows traffic routed via Telstra Global (202.84.x.x) after ~19ms, indicating entry at a distant POP instead of an EU POP
- Tested multiple proxy hostnames (sakura, shuttle, caboose, gondola, turntable) — all in 66.33.22.0/24, all ~180ms from this server
- The same TCP proxy responds in ~60ms from other European ISPs
- Your EU edge itself is functioning normally; the issue is specific to AS24940 routing towards it
This matches the failure mode documented in your edge-networking docs: "all traffic consistently routes to a single, distant edge."
Request: Could your network team check the anycast announcement and BGP peering for 66.33.22.0/24 towards Hetzner/AS24940? The consistent misrouting suggests a peering or announcement issue specific to that ASN.
Affected service:
- Project: Footmex/prod (98eedb1e-22cb-46d0-b234-054e2b9ebb60)
- Service: Postgres-tNcb (96351671-72fc-4653-b581-ae80d2245e4e)
- Region: europe-west4-drams3a
- TCP proxy: sakura.proxy.rlwy.net:25896 → port 5432
Additional context: our Postgres service was auto-updated to PG18 around Aug 22 ~23:00 CET. Could you check whether that update changed anything about our TCP proxy assignment or routing? We want to rule out that the misrouting started with the update.
Follow-up with additional context from the Railway dashboard agent: the deployment history for Postgres-tNcb shows no deployments since July 21, and the config already shows image ghcr.io/railwayapp-templates/postgres-ssl:18. So the Aug 22-23 incident left no deployment record. To correlate with our error spike (Aug 23 00:00-01:00 UTC: transaction timeouts, connection resets on both our Postgres services), could you please check from your side:
Container restart/event logs for Postgres-tNcb on Aug 22-23
TCP proxy assignment history — was sakura.proxy.rlwy.net:25896 reassigned, or did its routing change on Aug 22-23?
Auto-update logs — exact timestamp of the PG18 image update (if any)
Any edge/BGP configuration changes affecting 66.33.22.0/24 towards AS24940 on or around Aug 22-23
Items 2 and 4 are the most important: our per-query latency evidence suggests either a long-standing misroute or one introduced around that window.
Happy to provide full traceroute/MTR output or run additional tests from our server on request.
2 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
Thanks for confirming the root cause. Due to production urgency we've migrated the database off Railway to a local instance, so we're no longer affected because we are no longer the customer. We will probably migrate our other databases and close our railway, most probably will not use here anymore, this was critical issue and we can not risk our production database to that kind of issiues. Still, it would be great if the network team fixes the AS24940 routing — we may return in the future. You can close the ticket.
a month ago
I want to correct the misinformation previously shared by a community member. We don't use Anycast for the TCP proxies, so Hetzner would need to help you get to the bottom of this.
Status changed to Awaiting User Response Railway • about 1 month ago
Status changed to Awaiting Conductor Response Railway • about 1 month ago