20 days ago
Project/Environment:
Project ID: 2beb8093-b7b7-40fc-98a4-54a9a8c8f837 (project name "harmonious-solace")
Environment ID: fe2e2d69-66c6-4e15-90ed-0ca633ac7ce9 (environment "production")
Region: US East
Symptom:
An Envoy service (built from 6ixfalls/supabase, /envoy root) acting as an API gateway cannot open a TCP connection to a PostgREST service on the same project's private network (*.railway.internal), while three other backend services on the identical setup (GoTrue/auth, Storage, Postgres-Meta) are reached successfully by the same Envoy instance without any issue.
Every request to /rest/v1/* returns:
HTTP 503
upstream connect error or disconnect/reset before headers. reset reason: remote connection failure
Envoy's own debug-level logs show the exact underlying cause is a genuine kernel-level ECONNREFUSED on connect, not a timeout or DNS failure:
[debug][connection] connecting to [fd12:9bea:7cc6:1:...]:3000
[debug][connection] connection in progress
[debug][connection] delayed connect error: Connection refused
This refusal happens in well under 1ms after the connect attempt starts, which reads as a local rejection rather than a genuine round-trip to a remote peer over the WireGuard mesh.
What we've verified does NOT explain it:
Target is reachable — from a plain root shell inside the same Envoy container (Railway Console), a raw TCP connection to the exact same private IPv6 address and port (fd12:9bea:7cc6:1:...]:3000) succeeds instantly and returns a valid HTTP response from PostgREST. This works reliably and repeatedly, including immediately before/after an Envoy-initiated attempt fails.
Not a UID/permissions issue — Envoy's own process (PID 1) and the console shell both run as uid=0(root).
Not DNS/config — cds.yaml's cluster definition for the rest cluster has the correct host/port, matching what the shell successfully connects to. No dns_lookup_family, source_address, bind_config, or transport_socket overrides anywhere in the generated config, and no custom socket_interface (e.g. sockmap) configured in the Envoy bootstrap.
Not IPv4/IPv6 family selection — the target host resolves to both an IPv4 and IPv6 private address; both were individually confirmed reachable via raw shell TCP tests. Envoy only ever attempts the IPv6 address (per its own logs), same as the working auth cluster, which also connects over an IPv6 private address — so this doesn't look like a Happy-Eyeballs-related failure either.
Not the "Enable Outbound IPv6" per-service toggle — tested both on and off on the PostgREST service; no change in behavior.
Not region/network allocation on the PostgREST service — reassigned its region to force a fresh network allocation; no change.
Not a stale/corrupted service registration — we created a brand-new PostgREST service from scratch (new service, new private hostname postgrest-9b20.railway.internal, new IP, same source repo/root dir /rest, same environment variables). Envoy was pointed at this new service. Identical failure — same Connection refused from Envoy, same instant success from a raw shell TCP test to the same address.
Isolation: This appears specific to Envoy connecting to a PostgREST (Haskell/Warp-based) service specifically — two independent PostgREST instances both fail this way through Envoy, while Node.js-based (Storage, Postgres-Meta) and Go-based (GoTrue) services on the identical private network, reached by the identical Envoy instance, all work with zero issues.
Question: Is there any per-service or per-process network policy on Railway's private mesh (conntrack/eBPF-based egress filtering, connection rate limiting, etc.) that could apply differently to a proxy's outbound connection pool (Envoy, opening a new connection every 5s via active health checking) versus a single ad-hoc shell-initiated connection — and that might specifically be tripped by a Warp/Haskell-based listener's socket behavior? Any known issue connecting Envoy-based services to PostgREST over private networking?
Happy to provide full deployment IDs, additional debug logs, or a live repro window if useful.
2 Replies
20 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 20 days ago
Status changed to Solved viniciuskingjoe • 20 days ago
20 days ago
Marked as solved by accident, this isn't actually resolved — please reopen
Status changed to Open Railway • 20 days ago
20 days ago
Found it — PGRST_SERVER_HOST=* (dual-stack) has some bug in this environment; PGRST_SERVER_HOST=*6 (IPv6-only) fixes it immediately
Status changed to Solved viniciuskingjoe • 20 days ago