Consistently slow Postgres connection/query latency (2-7s for SELECT 1) — seeking origin-side diagnosis
daevid-thegreat
HOBBYOP

2 months ago

Hi Railway team,

I'm running a Postgres service (project: postaearn, service ID visible in my dashboard, host hayabusa.proxy.rlwy.net:54412) and seeing consistent, high latency on even the simplest possible query, from multiple independent test angles. I'd like help understanding whether this is expected for my current plan/region, or whether something's wrong on the origin side.

What I've measured, from my own machine (not through any app):

Raw TCP connect to the proxy host: ~0.47s (fast, as expected)

ICMP ping to the same host: 200-345ms, with 46.7% packet loss across a 15-packet sample

A full psql session running nothing but SELECT 1 (connect + auth + query + response): consistently 2-7 seconds across repeated runs, both with SSL enabled and with sslmode=disable (ruling out certificate/TLS negotiation as the cause)

DNS resolution for the proxy hostname: ~180-469ms for a single query (also unusually slow)

Why I don't think this is my network: a plain curl to https://www.google.com from the same machine, at the same time, completed in 0.82s. General connectivity is healthy; the slowness is specific to this Postgres connection.

Why I don't think this is distance/region alone: I ran the identical SELECT 1 test against a completely different provider (Neon) from the same machine and got similarly slow, similarly variable results (multiple seconds, inconsistent across runs) — which suggests something in common between these tests (possibly my own path to US-hosted database infrastructure generally, or something else) rather than an issue specific to Railway's proxy. I'm raising it with you first since you have visibility into origin-side timing that I don't.

What I'd like to know:

Does Railway have server-side timing/logs for connections to this service that could show where the time is actually going (proxy handshake, auth, query execution)?

Is there anything about the public proxy specifically (vs. a hypothetical direct connection) that would explain multi-second overhead on a trivial query?

Is this latency level expected/typical for my current plan, or does it indicate a problem?

Happy to provide more diagnostics (timestamps, additional time psql runs) if useful. This is affecting a production application currently serving real users, so I'd appreciate a look when possible.

Thanks,

Awaiting Conductor Response$10 Bounty

1 Replies

Railway
BOT

2 months ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 2 months ago


testuchunmaxsus
PRO

2 months ago

Your Neon control test already contains the answer: the cost is round trips on your path, not the origin. Here is the arithmetic, using your own numbers.

A fresh psql -c "SELECT 1" is not one round trip — it is about eight: DNS lookup, TCP handshake, TLS negotiation (1–2 RT), startup packet to auth request, SCRAM-SHA-256 client-first to server-first, SCRAM client-final to server-final, query to ReadyForQuery, terminate.

At your measured 200–345 ms RTT that is 1.6–2.8 s before Postgres executes anything. That is the floor of your 2–7 s.

The rest is the packet loss. If ~47% loss is real at the TCP layer (test 3 below), every lost segment costs a full retransmission timeout, and the initial RTO is 1 second (RFC 6298). Across ~8 exchanges, 1–3 stalls is the most likely outcome: P(0 lost) is about 0.7%, P(1) about 5%, P(2) about 14%, P(3) about 25%. So 1.6–2.8 s of base round trips plus 1–3 s of RTO stalls equals 2–7 s, and the run-to-run variance is simply the loss being random. That matches your observation precisely.

Why sslmode=disable did not help: it removes 2 of the ~8 round trips. On a healthy path that is noticeable; on a lossy 300 ms path it is noise. Correctly ruling out TLS does not rule out the network.

Why the curl to google.com is misleading: google.com is anycast, so you reach a POP inside or adjacent to your own ISP and never touch the long-haul path your database traffic takes. (0.82 s to a CDN edge is itself on the slow side.) Your Neon test is a valid control, and it agrees with everything above.

Three tests that settle it, none of which need Railway-side access.

  1. Separate connection cost from query cost. Open one psql session and stay in it, run \timing on, then run SELECT 1; ten times. Steady state should be about 1 RTT (200–350 ms). If it is, connection establishment is 100% of the problem and the server is fine.

  2. Confirm the server is idle-fast. In that same session run EXPLAIN (ANALYZE) SELECT 1; and execution time will be far below 1 ms. Wall clock minus that is network, by definition. This is the origin-side timing you asked for, and you can produce it yourself.

  3. Check whether the loss is real. Routers rate-limit ICMP, so 46.7% ping loss is frequently a false alarm. Measure at TCP instead: mtr -T -P 54412 -c 100 hayabusa.proxy.rlwy.net . If loss starts at a middle hop and persists through the final hop, the path is genuinely lossy; if only intermediate hops show loss while the final hop is clean, disregard it.

Your three questions, directly.

Server-side timing logs? Test 2 gives you the equivalent. Nothing on the origin side will show multi-second execution for SELECT 1.

Does the public proxy add multi-second overhead? No. It is a TCP passthrough inside the datacenter, sub-millisecond. What it adds is distance — you cross the public internet on every connection. Private networking skips both the proxy and the internet.

Is this expected for your plan? Plan tier does not affect the network path. Hobby and Pro use the same proxy.

What to change, in order of impact.

  1. Keep production traffic off this path entirely. If your app runs on Railway in the same project and environment, connect via the Postgres DATABASE_URL reference variable, which resolves to postgres.railway.internal:5432. Same datacenter, no proxy, no internet hop, and SELECT 1 lands well under a millisecond. If your production app is currently dialing hayabusa.proxy.rlwy.net, that single change is the fix.

  2. If the app must live outside Railway, never open a connection per request. Use a pool that holds connections open — asyncpg.create_pool, SQLAlchemy QueuePool, or pgbouncer. You pay the ~8 round trips once at startup instead of on every query, which turns 2–7 s into roughly 1 RTT.

  3. Move the database closer. Railway lets you choose the region per service. If you and your users are not in the US, redeploying Postgres to the nearest region reduces every RTT above proportionally.

  4. Fix DNS separately. 180–469 ms for a single lookup means a slow or distant resolver, so point at 1.1.1.1 or 8.8.8.8. A long-lived app caches this, so it mostly distorts your CLI measurements.

Short version: the database is healthy, the proxy is not adding seconds, and your own Neon result already said so. You are measuring connection establishment across a lossy long-haul path. Private networking, or a persistent pool, removes it.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...