a month ago
After converting my standalone MySQL to HA MySQL (3 data nodes + HAProxy, mysql:9.4, MySQL 9.4.0), every query routed through the HAProxy endpoint is ~4× slower than talking to a MySQL node directly — despite all services being in the same region. This adds up badly on query-heavy pages (a ~40-query admin page went from ~100 ms of DB time to ~450 ms+).
Measurements (run from inside my app container limousine-booking-stable, over the private network, SELECT 1 × 30–50 samples on a single already-open PDO connection):
Target Per-query latency
mysql-ha.railway.internal (HAProxy) ~12–15 ms
mysql-1.railway.internal (primary node, direct) ~2.5–5 ms
mysql-3.railway.internal (direct) ~1–4 ms
HAProxy IP 10.138.217.204 (direct) ~18 ms
HAProxy IP 10.175.205.126 (direct) ~17 ms
→ ~9–11 ms of per-query overhead attributable to the HAProxy layer.
What I've already ruled out
Not region placement: app, HAProxy, and primary all report RAILWAY_REPLICA_REGION=asia-southeast1-eqsg3a.
Not a single bad instance: mysql-ha resolves to both HAProxy IPs and both measure ~17–18 ms — the overhead is uniform, not one distant replica.
Not the DB itself: innodb_buffer_pool_size = 15 GB, buffer-pool hit ratio 99.9%, all 3 group-replication members ONLINE; direct-to-node queries are fast (~3 ms).
Hypothesis
This looks like a proxy-layer issue rather than network distance — possibly Nagle's algorithm / missing TCP_NODELAY on the HAProxy front-end or back-end (delayed-ACK interaction with MySQL's small request/response packets), or HAProxy instance resourcing.
Questions / ask
Is TCP_NODELAY set on the managed HAProxy's frontend and backend connections?
Can the HAProxy per-query latency be tuned/reduced toward the ~3 ms direct-node figure? Would resizing the HAProxy instances help?
Is ~10 ms/query in-region overhead expected for HA MySQL, or is something misconfigured in my cluster?
Happy to provide the exact benchmark script or run further diagnostics. Thanks!
3 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 29 days ago
a month ago
Your measurements strongly suggest a fixed per-round-trip latency in the proxy path rather than MySQL execution time or instance placement.
The most useful next step is to distinguish “proxy overhead per query” from “proxy overhead per TCP round trip”.
Since you're already using a single persistent PDO connection, I'd run this experiment:
- Benchmark 50 sequential
SELECT 1queries. - Benchmark the same logical workload as a multi-statement query / batched request where possible.
- Compare p50/p95/p99, not just averages.
For example, if you observe something like:
- 1 query through HAProxy: ~13 ms
- 10 sequential queries: ~130 ms
- 10 queries in one round trip: ~15–20 ms
then the evidence points strongly toward a per-request transport/proxy latency rather than MySQL execution or HAProxy throughput.
I'd also test whether the ~10 ms gap is independent of query size:
SELECT 1- small indexed lookup
- larger result set
A roughly constant delta would indicate fixed round-trip latency. A delta that grows with payload would instead suggest buffering or throughput effects.
The fact that both HAProxy instances show similar latency and all services report the same Railway region makes a geographic placement issue unlikely. Likewise, the persistent connection benchmark rules out connection establishment/TLS handshake costs.
I wouldn't assume TCP_NODELAY is the root cause without packet timing, though. There are two TCP legs:
App → HAProxy
HAProxy → MySQL
So a delayed-ACK/Nagle interaction could theoretically occur on either leg. A packet capture or HAProxy timing metrics would be the cleanest way to identify where the ~10 ms gap is introduced.
One particularly useful diagnostic would be timestamping:
App sends COM_QUERY
→ HAProxy forwards to backend
→ MySQL responds
→ HAProxy forwards response
That would show whether the delay occurs before forwarding, on the backend leg, or during response delivery.
If Railway confirms this is expected proxy-path latency, I'd consider exposing a lower-latency HA endpoint or documenting that HAProxy adds a fixed ~10 ms/query cost. At 40 sequential queries/page, that's effectively ~400 ms of application-visible latency even when the database itself is healthy.
As an immediate application-side mitigation, reducing round trips (batching, eager loading, eliminating N+1 queries) will have an outsized impact compared with optimizing query execution time.
a month ago
Ran your experiment — the evidence is conclusive:
- It's per-round-trip, not per-query execution.
10× SELECT 1 sequential through HAProxy: 58.3 ms (~5.8 ms/RT)
Same 10 statements batched into one round trip: 6.7 ms — i.e. a single round trip's cost.
Collapsing 10 round trips → 1 takes it from 58 ms to 6.7 ms. Direct-to-node shows the same shape (11.2 ms → 1.1 ms). So the cost is per round trip, not query execution or proxy throughput.
- The delta is a fixed round-trip cost, independent of payload.
Delta HAProxy−direct: SELECT 1 +4.5 ms, indexed 1-row +5.6 ms, 500-row result +2.3 ms.
The delta stays flat — and actually shrinks as a fraction on the 500-row query (which is transfer-dominated and nearly identical on both paths). That rules out buffering/throughput effects and points to a fixed per-round-trip transport latency in the proxy path.
- Tail latency is worse through the proxy too: SELECT 1 p99 17.5 ms via HAProxy vs 3.7 ms direct. And across runs the absolute overhead varies (~5 ms now, ~10–13 ms earlier), which suggests it's sensitive to proxy-path scheduling/load rather than being a static network distance.
Agreed that pinning down which TCP leg (app→HAProxy vs HAProxy→MySQL) needs a packet capture / HAProxy timing on the managed side — I don't have access into the HAProxy container to capture that. If Railway can share HAProxy timing metrics or confirm whether this fixed per-RT cost is expected, that'd close it out. In the meantime I'm reducing round trips app-side (caching reference data, eager-loading, killing N+1), which — at ~40 sequential queries/page — is the highest-leverage mitigation, as you said.
Thanks for the sharp analysis on this.
8 days ago
Before changing clients, I’d check which MySQL driver PDO is using. PHP’s mysqlnd already enables TCP_NODELAY on TCP connections—it explicitly calls setsockopt(..., TCP_NODELAY, 1). So the suggestion that PDO leaves Nagle enabled doesn’t apply generally. PHP source
You can check from your existing benchmark connection:
echo $pdo->getAttribute(PDO::ATTR_CLIENT_VERSION), PHP_EOL;If that reports mysqlnd, switching to Python solely to enable TCP_NODELAY is unlikely to address the missing setting being suggested.
There’s also one useful control to add to your existing benchmark: verify that the proxied and direct connections reach the same MySQL member. Run this once on each connection, outside the timed loop:
SELECT @@hostname, @@server_uuid;Then compare two already-open connections to that same member, alternating direct/proxied measurements. Since your overhead changed between runs, alternating them helps reduce the influence of changing background load.
For the managed-side investigation, I’d make the request more specific than “HAProxy timing metrics”:
- Kernel TCP RTT for both legs: HAProxy exposes
fc_rtt(us)for the client connection andbc_rtt(us)for the backend connection, where supported. - Retransmissions and proxy CPU throttling/scheduling during the same benchmark window.
Those RTT measurements could show whether one transport leg accounts for much of the delay. They aren’t per-query timings, but they help decide where to investigate next. HAProxy TCP metrics
A distinction that matters here: ordinary HAProxy TCP logs describe the connection/session, not each SQL statement within your persistent PDO connection. A low backend connection time (Tc) would therefore not explain or rule out the delay on subsequent queries. HAProxy TCP logging
I’d base any resizing decision on evidence of CPU contention or throttling. Your measurements establish extra latency along the proxied path, but don’t yet distinguish network transport from time spent waiting inside the proxy.