a month ago
Summary
Our production API experienced consistent timeouts on database-dependent endpoints when Railway routed requests through the Singapore region. The same endpoints worked normally on the US East replica.
Service: Ulink / ulink-api / production
Domain: https://api.ulink.ly
Affected endpoint example: GET /projects
Date observed: 2026-08-24
Impact
Users in Indonesia, or users connecting through an Indonesian VPN, experienced request timeouts after approximately 15–35 seconds. This prevented them from loading projects and other database-dependent resources.
Non-database endpoints continued working.
Expected behavior
Authenticated GET /projects requests should return in under a few seconds, regardless of which Railway replica handles the request.
Observed behavior
When handled by Singapore:
- GET /health returned 200 normally.
- GET /debug/protected returned 200, confirming that the application and JWT verification were working.
- GET /projects produced no response body and eventually timed out.
- Railway recorded some aborted requests as HTTP 499 with txBytes=0.
- Direct tests waited for more than 30 seconds without receiving any bytes.
When handled by US East:
- GET /projects returned 200 or 304.
- Normal backend duration was approximately 175–420 ms.
Affected Railway deployment
Deployment ID:
1aba0385-4268-419c-bcae-1b282538116e
Singapore region:
asia-southeast1-eqsg3a
Affected Singapore instance:
2be43b2c-edc2-4cc6-aa33-861eb9fafe8e
Healthy US East instance:
5e926911-0cbd-4578-bd43-9a7e505a3d05
Database dependency
The endpoint connects to a Supabase PostgreSQL transaction pooler over TCP port 6543. Railway network flow logs from the Singapore replica showed repeated small TCP connection attempts and inbound TCP_INVALID_SYN drops involving these pooler addresses:
- 52.59.152.35
- 18.198.30.239
- 18.198.145.223
The US East replica showed normal multi-kilobyte PostgreSQL traffic to the same pooler.
Supabase database network restrictions were checked and allow both IPv4 and IPv6:
- 0.0.0.0/0
- ::/0
Workaround
We removed the Singapore replica and kept one US East replica. No application-code change was required.
Current healthy deployment:
Deployment ID:
8d1f4615-4db6-46c9-b887-0518da32e142
Region:
us-east4-eqdc4a
Instance:
be28ca87-b018-42a7-a71e-8d7ca7ac69e1
Application commit:
2ff3462afec2df9b7162d6b67d58d18576b0571d
Verification after workaround
At 2026-08-24 08:45 UTC:
- GET /health: 200 in 0.84 seconds
- GET /debug/protected: 200 in 0.82 seconds
- GET /projects: 200 in 1.19 seconds
- Railway upstream duration for /projects: 201 ms
- Edge region: us-east4-eqdc4a
Request for Railway support
Could you investigate outbound TCP/NAT or routing problems between Railway instances in asia-southeast1-eqsg3a and the Supabase PostgreSQL transaction pooler on port 6543?
We would also like to know whether TCP_INVALID_SYN drops in the Singapore instance’s flow logs indicate a known networking issue and whether it is safe to re-enable the Singapore replica.
Similar previously confirmed incident
This failure pattern appears similar to a previously reported Railway incident titled:
“Outbound networking failure - all external connections lost (March 15, 07:50–08:27 UTC)”
In that case, Railway confirmed a known outbound TCP egress issue affecting connections to the Supabase PostgreSQL pooler on port 6543.
Our incident occurred on a different date and appeared isolated to the Singapore region, but the symptoms are similar:
- Supabase pooler connections timed out.
- The application remained healthy.
- Another Railway region could reach the same database normally.
- Railway flow logs showed abnormal TCP behavior.
- Connectivity recovered after removing the affected replica without changing application code.
Could you confirm whether our Singapore replica was affected by the same known class of outbound TCP egress issue?
1 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
Same symptoms, Singapore → Supabase transaction pooler (6543). We diagnosed it in production and found a working fix — details below.
Setup: Railway Singapore (asia-southeast1-eqsg3a), TanStack Start app, postgres.js → Supabase transaction pooler on 6543 (Supabase project also in ap-southeast-1, so this wasn't even cross-region). A second service on the same Railway host (an MQTT ingest worker writing to the same database every few seconds) never had a single problem — that contrast turned out to be the key diagnostic.
Symptoms, matching this thread: after 30–60 minutes of quiet/bursty use, every DB-bound request hangs indefinitely. Non-database routes keep answering in ~100ms, health checks that don't touch the DB stay green, deploys succeed. Cloudflare in front eventually surfaces it as 524/504. A container restart always fixes it — until the next time.
What we found on the database side (worth checking if you're hitting this — it's the smoking gun): during a hang, pg_stat_activity showed our sessions in state=active, wait_event=ClientRead for 60+ seconds — Postgres holding a query open, waiting for a client that was gone. The pooler was never full (we peaked around half the backend limit). So this is not pool exhaustion: something on the path silently kills quiet TCP flows, and the client never learns. The steady-traffic worker survives because its connections never sit quiet; the bursty web app's do.
Amplifier: postgres.js has no pool-acquire timeout and no per-query timeout, so a query dispatched onto a dead socket waits until TCP retransmission gives up (minutes). A dropped connection therefore presents as a total silent outage rather than a burst of errors. Tuning that helped bound it: keep_alive: 10, idle_timeout: 30, max_lifetime: 600, plus an onclose log so drops are visible at all. (Warning: do not set connection: { statement_timeout } in postgres.js against the transaction pooler — it's sent as a startup parameter and Supavisor rejects it, killing every connection.)
The fix that worked: bypass the pooler entirely. Supabase's direct endpoint (db..supabase.co:5432) is IPv6-only. Railway containers have no IPv6 route by default (ENETUNREACH), but there's a per-service setting — Settings → "Outbound IPv6" — that enables it. Toggle that on, point DATABASE_URL at the direct host on 5432 (user postgres, sslmode=require), redeploy. Since switching: zero hangs under exactly the usage that reliably reproduced them, and noticeably lower per-query latency. Our worker is still on the pooler deliberately, as a control.
Honest caveats: this is hours of clean running, not weeks; and it doesn't establish whose component drops the flows (Railway egress/NAT, the AWS LB in front of Supavisor, or in between) — only that removing the pooler hop removes the problem. Direct connections also bypass Supavisor's multiplexing, so this suits apps with a few long-lived processes and small pools; serverless-scale fan-out still needs the pooler (or Supabase's dedicated IPv4 add-on for the direct endpoint if you can't enable IPv6).