2 months ago
Project: supportive-compassion (ID d108e9e2-ad56-4f23-a3dd-050dd103bbaa), deployed from the official railway.com/deploy/supabase template (6ixfalls/supabase-postgres, PostgREST v14.12, Envoy).
Symptom: Every request to /rest/v1/* through the public Envoy gateway returns:
HTTP 503
upstream connect error or disconnect/reset before headers. reset reason: remote connection failure
/auth/v1/* (GoTrue) requests through the same Envoy instance work fine, so this looks specific to the Postgrest route/cluster, not Envoy as a whole.
What I've checked/tried, none of which changed anything:
- Postgrest's own deploy logs are clean on every restart -- connects to Postgres successfully, schema cache loads fine (7 relations), no errors.
-
- Envoy's REST_HOST variable correctly points to postgrest.railway.internal, matching Postgrest's actual private domain.
-
- Redeployed Postgrest (x2), Envoy (x2), Postgres, and Supavisor -- identical error every time, byte-for-byte.
-
- No incident showing on status.railway.com or in the community forum matching this.
Possibly relevant: the issue first appeared shortly after I redeployed one unrelated service (GoTrue Auth, for a config change) and, separately, an accidental UI click briefly created a phantom service on the project canvas that I deleted before it ever deployed. Around that same window, the Railway dashboard itself showed me a "Networking info temporarily unavailable -- Domains and TCP proxy details could not be loaded" toast. I can't confirm causation, but the timing is suspicious enough to mention.
This looks like a private-networking issue between two services in the same project/environment, not an app-level misconfiguration. Any help figuring out what's actually failing on the backend would be appreciated -- happy to provide deployment IDs or logs on request.
4 Replies
2 months ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 2 months ago
2 months ago
Thanks for the detailed explanation -- that theory lines up well with what I'm seeing. However: PGRST_SERVER_HOST was already set to * (not the default !4) before I saw this reply, so that wasn't the gap. To be thorough I still redeployed both Postgrest and Envoy fresh just now, in that order, and retested -- identical result, byte-for-byte:
HTTP 503
upstream connect error or disconnect/reset before headers. reset reason: remote connection failure
One more data point that might help: around 2026-08-14 07:17-07:21 PDT, Postgrest's own logs briefly showed it losing its connection to Postgres too (postgres.railway.internal), and the retry log line shows it attempting both an IPv4 address (10.179.238.132) and an IPv6 address (fd12:2ab2:6b3c:1:5000:4a:3833:ee84) for that host, both timing out simultaneously before it self-recovered a few seconds later. So dual-stack addresses are definitely in play internally, and there was a real (if brief) connectivity blip around that time -- but Envoy->Postgrest has been consistently failing before, during, and after that window, not just then, so it doesn't look purely transient either. Happy to try anything else you'd suggest.
2 months ago
Thanks for the detailed steps — I ran them from Envoy's console and got a very clean reproduction, and it points at something more specific than a generic private-networking blip.
getent hosts comparison (as suggested):
postgrest.railway.internal -> fd12:2ab2:6b3c:1:d000:11f:3d96:b8b7 (IPv6 only)
gotrue-auth.railway.internal -> fd12:2ab2:6b3c:1:d000:66:af30:3bb4 (IPv6 only)
Both resolve IPv6-only via getent hosts. getent ahosts (full dual-stack lookup) shows both also have an IPv4 address available (10.150.184.183 for postgrest, 10.176.59.180 for gotrue-auth), so DNS itself isn't the differentiator.
Direct connect test to the exact IPv6 address Envoy's cluster resolved:
$ timeout 3 bash -c '</dev/tcp/fd12:2ab2:6b3c:1:d000:11f:3d96:b8b7/3000'
bash: connect: Connection refused
Same test to PostgREST's IPv4 address (10.150.184.183:3000) connects immediately. So this isn't a routing/reachability failure — it's an active refusal on the IPv6 address specifically.
Envoy's own admin stats (127.0.0.1:9901/clusters) confirm this is the whole story:
rest::[fd12:...:b8b7]:3000::health_flags::/failed_active_hc
rest::[fd12:...:b8b7]:3000::cx_connect_fail::9
auth::[fd12:...:3bb4]:9999::health_flags::healthy
The rest cluster's only endpoint is permanently failed_active_hc because Envoy can't open a TCP connection to it over IPv6 — that's the entire cause of the 503s. auth's endpoint is healthy on the exact same kind of IPv6 address, so GoTrue is reachable over IPv6 while PostgREST is not.
I checked PGRST_SERVER_HOST and tried the real fix: it's set to *, which — per PostgREST's own docs — actually means "IPv4 only," not "all interfaces." I changed it to !4 (PostgREST's documented syntax for "any host, IPv4 and IPv6") and redeployed PostgREST. The deploy picked up the new value, but PostgREST's own startup log still says:
API server listening on 0.0.0.0:3000
— the IPv4-only wildcard, not a dual-stack bind. A fresh direct connect test to PostgREST's new IPv6 address after that redeploy still got Connection refused.
So even a correctly configured dual-stack listen (!4) isn't resulting in PostgREST actually accepting IPv6 connections in this container. That points at IPv6 not being usable inside the PostgREST service's own container/network namespace, even though Railway's private-networking layer is still handing out an IPv6 address for it that other services (Envoy) try to connect to. I've reverted PGRST_SERVER_HOST back to * for now since !4 had no effect and I didn't want to leave a non-default value in place without a reason.
Given that, this looks like it needs a look from your side at whether IPv6 is actually enabled in the PostgREST service's container/network namespace for this project, or whether Envoy's DNS resolution for this specific private hostname should be pinned to IPv4-only until that's sorted out. Happy to run any other diagnostics from the Envoy or PostgREST console if useful.
2 months ago
How about set on the PostgREST service:
PGRST_SERVER_HOST=*6
(:: works too; *6 is the documented reserved form and falls back to IPv4 gracefully.) Redeploy PostgREST, then confirm the startup log reads API server listening on :::3000 or [::]:3000 rather than 0.0.0.0:3000. Then re-run the same </dev/tcp/fd12:…/3000 test from Envoy's console — it should connect — and Envoy's rest cluster health flag should flip to healthy within a health-check interval or two.
If the log still says 0.0.0.0 after that, the variable isn't reaching the process rather than IPv6 being broken. Check with env | grep PGRST in the PostgREST console, and check whether the template's start command passes --server-host or mounts a config file — an explicit CLI flag or config file beats the env var.
lockerbill
How about set on the PostgREST service: PGRST_SERVER_HOST=*6 (:: works too; *6 is the documented reserved form and falls back to IPv4 gracefully.) Redeploy PostgREST, then confirm the startup log reads API server listening on :::3000 or [::]:3000 rather than 0.0.0.0:3000. Then re-run the same </dev/tcp/fd12:…/3000 test from Envoy's console — it should connect — and Envoy's rest cluster health flag should flip to healthy within a health-check interval or two. If the log still says 0.0.0.0 after that, the variable isn't reaching the process rather than IPv6 being broken. Check with env | grep PGRST in the PostgREST console, and check whether the template's start command passes --server-host or mounts a config file — an explicit CLI flag or config file beats the env var.
2 months ago
Your suggestion was the fix: setting
PGRST_SERVER_HOST=*6 got Postgrest listening on [::]:3000 instead of 0.0.0.0:3000, and once it
had a real IPv6 dual-stack bind, Envoy's rest cluster started passing health checks and
/rest/v1/* came back online. Production has been stable since.
For anyone hitting this later: the earlier !4 value (PostgREST's documented dual-stack syntax)
did NOT produce dual-stack binding in this container — the log still showed IPv4-only even after
redeploying. *6 was what actually worked. Thanks for the help tracking this down.