a month ago
Since the scheduled security patch (CVE-2026-15741 and others, applied ~Aug 21
per the in-app notification), connections to our Postgres instance via the
public TCP proxy (tramway.proxy.rlwy.net) have become significantly slower —
both from our local dev environment and from Prisma Studio connecting
directly to the production database.
Observed:
-
A bare
SELECT 1via psql over the public proxy takes ~2.5s round-trip(previously near-instant), suggesting the overhead is in connection
setup/auth, not query execution.
-
pg_stat_activityshows all idle connections are under ~8 minutes old —connections aren't persisting between requests the way they used to.
-
pg_postmaster_start_time()confirms the Postgres engine restarted on2026-08-21, matching the patch window.
-
Our production API (which reaches Postgres over Railway's private network,
not this public proxy) is unaffected — only connections originating from
outside the private network are slow.
This points at something changed in how the public proxy handles new
connections (e.g. auth negotiation, TLS, or routing) rather than the
Postgres engine itself, since CVE-2026-15741 is a narrow SQL-injection fix
in EXTRACT() deparsing and wouldn't be expected to add connection-level
latency.
Could you confirm whether the proxy/connection layer was also changed as
part of this patch window, and whether this latency is expected to resolve
on its own or needs a fix on our end?
Project: Operazor
Public proxy host: tramway.proxy.rlwy.net
2 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
You've got two symptoms here, and I think they're being treated as one: the ~2.5s connection time, and the fact that no connection is older than ~8 minutes. Slower auth wouldn't explain connections getting recycled at 8 minutes. An idle timeout would explain both: connections get reaped, so the client reconnects constantly and pays the full setup cost every time. The proxy was probably always that slow to connect; you just weren't paying that cost per request before. So the regression is more likely connection reuse dying, not necessarily the proxy itself becoming slower.
Two things I'd get you to run:
dig +short AAAA tramway.proxy.rlwy.netRailway's private network is IPv6-only, and if the proxy host picked up an AAAA record while your local IPv6 is broken, Happy Eyeballs could stall for ~2s before falling back to IPv4. That fits the constant delay, appearing overnight, external-only behavior, while queries themselves remain fast. You can confirm by pinning ?hostaddr=<the A record> in the connection URL.
Also, pg_stat_activity should be checked grouped by client_addr rather than in aggregate. Your production API is on the private network, so those connections should have an fd12:: address. If those are also capped at ~8 minutes, it's likely server-side — check idle_session_timeout. If only the proxy connections are young, then the proxy is likely reaping idle TCP connections, and:
?keepalives=1&keepalives_idle=120may help.
a month ago
Thanks for the reply and some follow-up with some more isolated data:
-
IPv6/Happy Eyeballs isn't it: tramway.proxy.rlwy.net has no AAAA record (checked via two resolvers), so there's no dual-stack race on this host.
-
It's not idle-connection reaping either, at least not primarily: idle_session_timeout and idle_in_transaction_session_timeout are both 0 on the Postgres side, and I isolated query cost from connection cost directly — opened one psql session, kept it open, and ran several SELECT 1 back to back on the same, already-established connection. Each one still took ~330ms round trip. So it's not reconnect overhead — a warm connection pays this cost per query too.
-
This isn't specific to one database: I get the same ~330ms per-query round trip on both Postgres services in the project (different hosts/ports), so it's not a single DB instance issue.
-
Raw ICMP ping to the proxy IP itself is ~175ms (vs ~13ms to 1.1.1.1 as a sanity baseline), which is roughly half the query RTT and consistent with a two-hop path: me → edge proxy (~175ms) → an internal hop to the actual Postgres pod (in europe-west4, where both DBs are hosted).