a month ago
Subject: Same-region PostgreSQL public-proxy latency discrepancy;
growing in-memory persistence backlog — preserve existing sessions
Our Railway application and PostgreSQL services are reported in us-west2.
The application connects through the Railway public proxy.
Application project:
b6a69107-64ed-4ad6-81e1-5550336a9e94
Application service:
cfbb87b4-e37d-4b81-8e9f-b2a73739d2fe
PostgreSQL project:
62c8662a-93eb-4a35-bf6e-53deacdce6b5
PostgreSQL service:
948e2e5b-16af-48ae-8c32-61aab4e9eb38
On 7 September 2026, read-only tests from the application container showed:
-
12:58 UTC: SELECT 1 approximately 4.5–5.2 ms.
-
13:04 UTC, separate connection: simple/parameterized SELECTs
approximately 137–140 ms.
These helper measurements do not establish the existing writer's RTT.
Observed public-writer sessions had no blocking PID and frequently showed
idle-in-transaction / ClientRead. Our serial per-row persistence amplifies
round-trip overhead; a batching fix is prepared but not yet deployed.
Between approximately 12:35 and 13:17 UTC, requested revisions increased
150→166 while confirmed revisions increased86→87; pending work grew64→79.
CRITICAL: outstanding snapshot payloads are RAM-backed. Please investigate
routing/proxy/connection latency WITHOUT restarting either service,
disconnecting existing sessions, terminating queries or changing networking.
Please identify any nondisruptive remedy, and explain the exact impact of
any proposed change before it is applied. No disruptive action is authorized
by this support request.
2 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 28 days ago
a month ago
Root cause: anycast edge routing on the public TCP proxy
You're connecting app→Postgres through the public TCP proxy (*.proxy.rlwy.net). That routes through Railway's anycast edge network — each TCP connection hits the nearest edge POP by network topology, not geography, then routes internally back to us-west2.
This is exactly why SELECT 1 varies 5–138 ms:
4.5–5.2 ms = connection hit a nearby US West edge POP, short internal hop
137–140 ms = connection hit a distant POP (EU/Asia), routed all the way back to us-west2
SELECT 1 is pure network round-trip — no disk I/O, no planning, no locks. 100% of that variance is network path, not Postgres. The query executes in <0.1 ms; the rest is proxy + transit.
This is documented behavior: Railway's edge docs confirm anycast routes "based on network topology, not geographic distance." Railway staff have directly acknowledged the public proxy adding "20–100ms impact" vs private networking.
Why the persistence backlog grows
Each 138 ms round-trip blocks the next query on that connection. With a bounded pool, queries queue behind slow round-trips, RAM-backed snapshots pile up, and the data-loss window grows. Root cause isn't disk or query slowness — your writes are taking the scenic route through a distant edge POP.
Nondisruptive fix: switch to private networking
This is an app env var change, not a Railway networking infrastructure change. You're not changing regions, proxy config, or TCP proxy settings — just which connection string your app uses:
Replace DATABASE_PUBLIC_URL with DATABASE_URL (private: postgresql://...@postgres.railway.internal:5432/railway)
Private networking uses encrypted WireGuard tunnels directly within Railway's infra — no edge, no anycast, no public internet hops
Expected latency: <1 ms consistently, same-region
No DB restart: new connections take the private path; old ones drain naturally
Bonus: zero egress costs
If the connection string can't change
Enable PgBouncer (Settings → PostgreSQL → Connection Pooling). Per Railway docs, "enabling or disabling pooling won't restart your database." Transaction mode multiplexes connections — a 138 ms spike on one doesn't block others.
Batch writes: group inserts into fewer transactions. Each transaction = one round-trip. 10 writes → 1 transaction cuts 10 anycast hops to 1.
synchronous_commit = local (nondisruptive, no restart):
ALTER SYSTEM SET synchronous_commit = local;
SELECT pg_reload_conf();
In single-region with no sync replicas, local gives the same durability — you just stop waiting for a remote WAL ack that doesn't exist.
Confirm the diagnosis (read-only, no changes)
SELECT state, wait_event_type, now()-query_start AS dur
FROM pg_stat_activity WHERE state IS NOT NULL ORDER BY query_start;
SHOW synchronous_commit;
Then run SELECT 1 from both the public proxy URL and postgres.railway.internal:5432. Private will be sub-ms; public will show the 5–138 ms variance.
Summary
Public proxy (now) Private networking (fix)
Path App → anycast edge POP → us-west2 App → WireGuard → us-west2
Latency 5–138 ms (variable) <1 ms (consistent)
Egress Billed Free
The 138 ms spikes aren't a Postgres, volume, or resource problem — they're the public proxy routing same-region traffic through a distant edge POP. Switch to DATABASE_URL and the variance disappears, with no DB restart, no session disconnects, no query kills, and no Railway networking changes.
samudraban
Root cause: anycast edge routing on the public TCP proxy You're connecting app→Postgres through the public TCP proxy (*.proxy.rlwy.net). That routes through Railway's anycast edge network — each TCP connection hits the nearest edge POP by network topology, not geography, then routes internally back to us-west2. This is exactly why SELECT 1 varies 5–138 ms: 4.5–5.2 ms = connection hit a nearby US West edge POP, short internal hop 137–140 ms = connection hit a distant POP (EU/Asia), routed all the way back to us-west2 SELECT 1 is pure network round-trip — no disk I/O, no planning, no locks. 100% of that variance is network path, not Postgres. The query executes in <0.1 ms; the rest is proxy + transit. This is documented behavior: Railway's edge docs confirm anycast routes "based on network topology, not geographic distance." Railway staff have directly acknowledged the public proxy adding "20–100ms impact" vs private networking. Why the persistence backlog grows Each 138 ms round-trip blocks the next query on that connection. With a bounded pool, queries queue behind slow round-trips, RAM-backed snapshots pile up, and the data-loss window grows. Root cause isn't disk or query slowness — your writes are taking the scenic route through a distant edge POP. Nondisruptive fix: switch to private networking This is an app env var change, not a Railway networking infrastructure change. You're not changing regions, proxy config, or TCP proxy settings — just which connection string your app uses: Replace DATABASE_PUBLIC_URL with DATABASE_URL (private: postgresql://...@postgres.railway.internal:5432/railway) Private networking uses encrypted WireGuard tunnels directly within Railway's infra — no edge, no anycast, no public internet hops Expected latency: <1 ms consistently, same-region No DB restart: new connections take the private path; old ones drain naturally Bonus: zero egress costs If the connection string can't change Enable PgBouncer (Settings → PostgreSQL → Connection Pooling). Per Railway docs, "enabling or disabling pooling won't restart your database." Transaction mode multiplexes connections — a 138 ms spike on one doesn't block others. Batch writes: group inserts into fewer transactions. Each transaction = one round-trip. 10 writes → 1 transaction cuts 10 anycast hops to 1. synchronous_commit = local (nondisruptive, no restart): ALTER SYSTEM SET synchronous_commit = local; SELECT pg_reload_conf(); In single-region with no sync replicas, local gives the same durability — you just stop waiting for a remote WAL ack that doesn't exist. Confirm the diagnosis (read-only, no changes) SELECT state, wait_event_type, now()-query_start AS dur FROM pg_stat_activity WHERE state IS NOT NULL ORDER BY query_start; SHOW synchronous_commit; Then run SELECT 1 from both the public proxy URL and postgres.railway.internal:5432. Private will be sub-ms; public will show the 5–138 ms variance. Summary Public proxy (now) Private networking (fix) Path App → anycast edge POP → us-west2 App → WireGuard → us-west2 Latency 5–138 ms (variable) <1 ms (consistent) Egress Billed Free The 138 ms spikes aren't a Postgres, volume, or resource problem — they're the public proxy routing same-region traffic through a distant edge POP. Switch to DATABASE_URL and the variance disappears, with no DB restart, no session disconnects, no query kills, and no Railway networking changes.
a month ago
Thanks—private networking and reducing round trips are useful directions.
Important topology detail: our app and PostgreSQL are in DIFFERENT Railway
projects, although both report us-west2. Railway's documentation scopes
private networking to the same project/environment.
Please clarify the supported private-connectivity path for this exact
topology and any migration/redeployment impact. A variable update must not
be treated as hot-migrating our existing Node process or PostgreSQL session.
Our old writer already uses one long transaction with separately awaited
INSERTs. Multi-row batching is prepared; putting statements in one
transaction alone does not make them one round trip.
Can Railway staff verify the actual public TCP-proxy path for the supplied
connection timestamps? The 5ms/138ms measurements do not independently
identify the edge POP or prove that all latency is network transit.
We are not authorizing networking, pooling or synchronous_commit changes
through this reply. Please distinguish a future configuration improvement
from a nondisruptive remedy for the currently established writer connection.