Same-region PostgreSQL proxy latency varies 5–138 ms; growing persistence backlog
mdfaridrokadiya-droid
PROOP

a month ago

Subject: Same-region PostgreSQL public-proxy latency discrepancy;

growing in-memory persistence backlog — preserve existing sessions

Our Railway application and PostgreSQL services are reported in us-west2.

The application connects through the Railway public proxy.

Application project:

b6a69107-64ed-4ad6-81e1-5550336a9e94

Application service:

cfbb87b4-e37d-4b81-8e9f-b2a73739d2fe

PostgreSQL project:

62c8662a-93eb-4a35-bf6e-53deacdce6b5

PostgreSQL service:

948e2e5b-16af-48ae-8c32-61aab4e9eb38

On 7 September 2026, read-only tests from the application container showed:

  • 12:58 UTC: SELECT 1 approximately 4.5–5.2 ms.

  • 13:04 UTC, separate connection: simple/parameterized SELECTs

    approximately 137–140 ms.

These helper measurements do not establish the existing writer's RTT.

Observed public-writer sessions had no blocking PID and frequently showed

idle-in-transaction / ClientRead. Our serial per-row persistence amplifies

round-trip overhead; a batching fix is prepared but not yet deployed.

Between approximately 12:35 and 13:17 UTC, requested revisions increased

150→166 while confirmed revisions increased86→87; pending work grew64→79.

CRITICAL: outstanding snapshot payloads are RAM-backed. Please investigate

routing/proxy/connection latency WITHOUT restarting either service,

disconnecting existing sessions, terminating queries or changing networking.

Please identify any nondisruptive remedy, and explain the exact impact of

any proposed change before it is applied. No disruptive action is authorized

by this support request.

$20 Bounty

2 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 28 days ago


samudraban
PRO

a month ago

Root cause: anycast edge routing on the public TCP proxy

You're connecting app→Postgres through the public TCP proxy (*.proxy.rlwy.net). That routes through Railway's anycast edge network — each TCP connection hits the nearest edge POP by network topology, not geography, then routes internally back to us-west2.

This is exactly why SELECT 1 varies 5–138 ms:

4.5–5.2 ms = connection hit a nearby US West edge POP, short internal hop

137–140 ms = connection hit a distant POP (EU/Asia), routed all the way back to us-west2

SELECT 1 is pure network round-trip — no disk I/O, no planning, no locks. 100% of that variance is network path, not Postgres. The query executes in <0.1 ms; the rest is proxy + transit.

This is documented behavior: Railway's edge docs confirm anycast routes "based on network topology, not geographic distance." Railway staff have directly acknowledged the public proxy adding "20–100ms impact" vs private networking.

Why the persistence backlog grows

Each 138 ms round-trip blocks the next query on that connection. With a bounded pool, queries queue behind slow round-trips, RAM-backed snapshots pile up, and the data-loss window grows. Root cause isn't disk or query slowness — your writes are taking the scenic route through a distant edge POP.

Nondisruptive fix: switch to private networking

This is an app env var change, not a Railway networking infrastructure change. You're not changing regions, proxy config, or TCP proxy settings — just which connection string your app uses:

Replace DATABASE_PUBLIC_URL with DATABASE_URL (private: postgresql://...@postgres.railway.internal:5432/railway)

Private networking uses encrypted WireGuard tunnels directly within Railway's infra — no edge, no anycast, no public internet hops

Expected latency: <1 ms consistently, same-region

No DB restart: new connections take the private path; old ones drain naturally

Bonus: zero egress costs

If the connection string can't change

Enable PgBouncer (Settings → PostgreSQL → Connection Pooling). Per Railway docs, "enabling or disabling pooling won't restart your database." Transaction mode multiplexes connections — a 138 ms spike on one doesn't block others.

Batch writes: group inserts into fewer transactions. Each transaction = one round-trip. 10 writes → 1 transaction cuts 10 anycast hops to 1.

synchronous_commit = local (nondisruptive, no restart):

ALTER SYSTEM SET synchronous_commit = local;

SELECT pg_reload_conf();

In single-region with no sync replicas, local gives the same durability — you just stop waiting for a remote WAL ack that doesn't exist.

Confirm the diagnosis (read-only, no changes)

SELECT state, wait_event_type, now()-query_start AS dur

FROM pg_stat_activity WHERE state IS NOT NULL ORDER BY query_start;

SHOW synchronous_commit;

Then run SELECT 1 from both the public proxy URL and postgres.railway.internal:5432. Private will be sub-ms; public will show the 5–138 ms variance.

Summary

Public proxy (now) Private networking (fix)

Path App → anycast edge POP → us-west2 App → WireGuard → us-west2

Latency 5–138 ms (variable) <1 ms (consistent)

Egress Billed Free

The 138 ms spikes aren't a Postgres, volume, or resource problem — they're the public proxy routing same-region traffic through a distant edge POP. Switch to DATABASE_URL and the variance disappears, with no DB restart, no session disconnects, no query kills, and no Railway networking changes.


samudraban

Root cause: anycast edge routing on the public TCP proxy You're connecting app→Postgres through the public TCP proxy (*.proxy.rlwy.net). That routes through Railway's anycast edge network — each TCP connection hits the nearest edge POP by network topology, not geography, then routes internally back to us-west2. This is exactly why SELECT 1 varies 5–138 ms: 4.5–5.2 ms = connection hit a nearby US West edge POP, short internal hop 137–140 ms = connection hit a distant POP (EU/Asia), routed all the way back to us-west2 SELECT 1 is pure network round-trip — no disk I/O, no planning, no locks. 100% of that variance is network path, not Postgres. The query executes in <0.1 ms; the rest is proxy + transit. This is documented behavior: Railway's edge docs confirm anycast routes "based on network topology, not geographic distance." Railway staff have directly acknowledged the public proxy adding "20–100ms impact" vs private networking. Why the persistence backlog grows Each 138 ms round-trip blocks the next query on that connection. With a bounded pool, queries queue behind slow round-trips, RAM-backed snapshots pile up, and the data-loss window grows. Root cause isn't disk or query slowness — your writes are taking the scenic route through a distant edge POP. Nondisruptive fix: switch to private networking This is an app env var change, not a Railway networking infrastructure change. You're not changing regions, proxy config, or TCP proxy settings — just which connection string your app uses: Replace DATABASE_PUBLIC_URL with DATABASE_URL (private: postgresql://...@postgres.railway.internal:5432/railway) Private networking uses encrypted WireGuard tunnels directly within Railway's infra — no edge, no anycast, no public internet hops Expected latency: <1 ms consistently, same-region No DB restart: new connections take the private path; old ones drain naturally Bonus: zero egress costs If the connection string can't change Enable PgBouncer (Settings → PostgreSQL → Connection Pooling). Per Railway docs, "enabling or disabling pooling won't restart your database." Transaction mode multiplexes connections — a 138 ms spike on one doesn't block others. Batch writes: group inserts into fewer transactions. Each transaction = one round-trip. 10 writes → 1 transaction cuts 10 anycast hops to 1. synchronous_commit = local (nondisruptive, no restart): ALTER SYSTEM SET synchronous_commit = local; SELECT pg_reload_conf(); In single-region with no sync replicas, local gives the same durability — you just stop waiting for a remote WAL ack that doesn't exist. Confirm the diagnosis (read-only, no changes) SELECT state, wait_event_type, now()-query_start AS dur FROM pg_stat_activity WHERE state IS NOT NULL ORDER BY query_start; SHOW synchronous_commit; Then run SELECT 1 from both the public proxy URL and postgres.railway.internal:5432. Private will be sub-ms; public will show the 5–138 ms variance. Summary Public proxy (now) Private networking (fix) Path App → anycast edge POP → us-west2 App → WireGuard → us-west2 Latency 5–138 ms (variable) <1 ms (consistent) Egress Billed Free The 138 ms spikes aren't a Postgres, volume, or resource problem — they're the public proxy routing same-region traffic through a distant edge POP. Switch to DATABASE_URL and the variance disappears, with no DB restart, no session disconnects, no query kills, and no Railway networking changes.

mdfaridrokadiya-droid
PROOP

a month ago

Thanks—private networking and reducing round trips are useful directions.

Important topology detail: our app and PostgreSQL are in DIFFERENT Railway

projects, although both report us-west2. Railway's documentation scopes

private networking to the same project/environment.

Please clarify the supported private-connectivity path for this exact

topology and any migration/redeployment impact. A variable update must not

be treated as hot-migrating our existing Node process or PostgreSQL session.

Our old writer already uses one long transaction with separately awaited

INSERTs. Multi-row batching is prepared; putting statements in one

transaction alone does not make them one round trip.

Can Railway staff verify the actual public TCP-proxy path for the supplied

connection timestamps? The 5ms/138ms measurements do not independently

identify the edge POP or prove that all latency is network transit.

We are not authorizing networking, pooling or synchronous_commit changes

through this reply. Please distinguish a future configuration improvement

from a nondisruptive remedy for the currently established writer connection.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...