Singapore replica cannot connect reliably to Supabase PostgreSQL pooler
mohn93
PROOP

a month ago

Summary

Our production API experienced consistent timeouts on database-dependent endpoints when Railway routed requests through the Singapore region. The same endpoints worked normally on the US East replica.

Service: Ulink / ulink-api / production

Domain: https://api.ulink.ly

Affected endpoint example: GET /projects

Date observed: 2026-08-24

Impact

Users in Indonesia, or users connecting through an Indonesian VPN, experienced request timeouts after approximately 15–35 seconds. This prevented them from loading projects and other database-dependent resources.

Non-database endpoints continued working.

Expected behavior

Authenticated GET /projects requests should return in under a few seconds, regardless of which Railway replica handles the request.

Observed behavior

When handled by Singapore:

  • GET /health returned 200 normally.
  • GET /debug/protected returned 200, confirming that the application and JWT verification were working.
  • GET /projects produced no response body and eventually timed out.
  • Railway recorded some aborted requests as HTTP 499 with txBytes=0.
  • Direct tests waited for more than 30 seconds without receiving any bytes.

When handled by US East:

  • GET /projects returned 200 or 304.
  • Normal backend duration was approximately 175–420 ms.

Affected Railway deployment

Deployment ID:

1aba0385-4268-419c-bcae-1b282538116e

Singapore region:

asia-southeast1-eqsg3a

Affected Singapore instance:

2be43b2c-edc2-4cc6-aa33-861eb9fafe8e

Healthy US East instance:

5e926911-0cbd-4578-bd43-9a7e505a3d05

Database dependency

The endpoint connects to a Supabase PostgreSQL transaction pooler over TCP port 6543. Railway network flow logs from the Singapore replica showed repeated small TCP connection attempts and inbound TCP_INVALID_SYN drops involving these pooler addresses:

  • 52.59.152.35
  • 18.198.30.239
  • 18.198.145.223

The US East replica showed normal multi-kilobyte PostgreSQL traffic to the same pooler.

Supabase database network restrictions were checked and allow both IPv4 and IPv6:

  • 0.0.0.0/0
  • ::/0

Workaround

We removed the Singapore replica and kept one US East replica. No application-code change was required.

Current healthy deployment:

Deployment ID:

8d1f4615-4db6-46c9-b887-0518da32e142

Region:

us-east4-eqdc4a

Instance:

be28ca87-b018-42a7-a71e-8d7ca7ac69e1

Application commit:

2ff3462afec2df9b7162d6b67d58d18576b0571d

Verification after workaround

At 2026-08-24 08:45 UTC:

  • GET /health: 200 in 0.84 seconds
  • GET /debug/protected: 200 in 0.82 seconds
  • GET /projects: 200 in 1.19 seconds
  • Railway upstream duration for /projects: 201 ms
  • Edge region: us-east4-eqdc4a

Request for Railway support

Could you investigate outbound TCP/NAT or routing problems between Railway instances in asia-southeast1-eqsg3a and the Supabase PostgreSQL transaction pooler on port 6543?

We would also like to know whether TCP_INVALID_SYN drops in the Singapore instance’s flow logs indicate a known networking issue and whether it is safe to re-enable the Singapore replica.

Similar previously confirmed incident

This failure pattern appears similar to a previously reported Railway incident titled:

“Outbound networking failure - all external connections lost (March 15, 07:50–08:27 UTC)”

In that case, Railway confirmed a known outbound TCP egress issue affecting connections to the Supabase PostgreSQL pooler on port 6543.

Our incident occurred on a different date and appeared isolated to the Singapore region, but the symptoms are similar:

  • Supabase pooler connections timed out.
  • The application remained healthy.
  • Another Railway region could reach the same database normally.
  • Railway flow logs showed abnormal TCP behavior.
  • Connectivity recovered after removing the affected replica without changing application code.

Could you confirm whether our Singapore replica was affected by the same known class of outbound TCP egress issue?

$20 Bounty

1 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


rhuddlestone
PRO

a month ago

Same symptoms, Singapore → Supabase transaction pooler (6543). We diagnosed it in production and found a working fix — details below.

Setup: Railway Singapore (asia-southeast1-eqsg3a), TanStack Start app, postgres.js → Supabase transaction pooler on 6543 (Supabase project also in ap-southeast-1, so this wasn't even cross-region). A second service on the same Railway host (an MQTT ingest worker writing to the same database every few seconds) never had a single problem — that contrast turned out to be the key diagnostic.

Symptoms, matching this thread: after 30–60 minutes of quiet/bursty use, every DB-bound request hangs indefinitely. Non-database routes keep answering in ~100ms, health checks that don't touch the DB stay green, deploys succeed. Cloudflare in front eventually surfaces it as 524/504. A container restart always fixes it — until the next time.

What we found on the database side (worth checking if you're hitting this — it's the smoking gun): during a hang, pg_stat_activity showed our sessions in state=active, wait_event=ClientRead for 60+ seconds — Postgres holding a query open, waiting for a client that was gone. The pooler was never full (we peaked around half the backend limit). So this is not pool exhaustion: something on the path silently kills quiet TCP flows, and the client never learns. The steady-traffic worker survives because its connections never sit quiet; the bursty web app's do.

Amplifier: postgres.js has no pool-acquire timeout and no per-query timeout, so a query dispatched onto a dead socket waits until TCP retransmission gives up (minutes). A dropped connection therefore presents as a total silent outage rather than a burst of errors. Tuning that helped bound it: keep_alive: 10, idle_timeout: 30, max_lifetime: 600, plus an onclose log so drops are visible at all. (Warning: do not set connection: { statement_timeout } in postgres.js against the transaction pooler — it's sent as a startup parameter and Supavisor rejects it, killing every connection.)

The fix that worked: bypass the pooler entirely. Supabase's direct endpoint (db..supabase.co:5432) is IPv6-only. Railway containers have no IPv6 route by default (ENETUNREACH), but there's a per-service setting — Settings → "Outbound IPv6" — that enables it. Toggle that on, point DATABASE_URL at the direct host on 5432 (user postgres, sslmode=require), redeploy. Since switching: zero hangs under exactly the usage that reliably reproduced them, and noticeably lower per-query latency. Our worker is still on the pooler deliberately, as a control.

Honest caveats: this is hours of clean running, not weeks; and it doesn't establish whose component drops the flows (Railway egress/NAT, the AWS LB in front of Supavisor, or in between) — only that removing the pooler hop removes the problem. Direct connections also bypass Supavisor's multiplexing, so this suits apps with a few long-lived processes and small pools; serverless-scale fan-out still needs the pooler (or Supabase's dedicated IPv4 add-on for the direct endpoint if you can't enable IPv6).


Welcome!

Sign in to your Railway account to join the conversation.

Loading...