a month ago
Two projects, production and staging. Same app, same commit, same env, both EU West.
Staging pages load instantly (~120ms). Production takes ~1 second.
It's the internal network. Measured from a shell inside each app container, median
of 10, on an already-open connection:
MySQL round trip production 9.4ms staging ~1ms
Redis round trip production 4.4ms staging ~1ms
My pages run 40-90 queries, so 9ms each becomes ~500ms a page.
It isn't the database doing work: SELECT 1 takes 9.32ms and a select from a
1.7M-row table takes 9.38ms. Both ends are idle - MySQL at 0.003 of 4 vCPU with a
100% buffer pool hit ratio, app container using 0.00 of 3 vCPU, nr_throttled 0.
Also, a CPU-only loop (no I/O) in the production container tracks the host load
average almost exactly, across four measurements on different days:
host load 42.6 -> loop 37ms
host load 10.5 -> loop 10ms
host load 7.0 -> loop 8ms
host load 6.0 -> loop 8ms
I was never throttled and using ~0.00 of my allocated CPUs in all four.
Why is the same round trip 9ms in one project and 1ms in the other, and can my
production services be moved?
2 Replies
a month ago
We've looked into this from our side and haven't found anything on the Railway platform that explains what you're seeing, so working it out means digging into your specific setup.
That's exactly what the Railway community is good at, so we'd like to open your thread as a community bounty. Railway pays a bounty to the community member who solves it, and threads like this usually get picked up quickly.
Opening it makes this entire thread public, including everything already posted. Nothing becomes public until you decide. Use the buttons below.
- Open to the community - Before you click, take a moment to edit or remove anything you'd rather not share. The thread becomes publicly visible right away.
- Keep it private and close the thread - Nothing becomes public and the thread closes.
Status changed to Awaiting User Response Railway • about 1 month ago
a month ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • about 1 month ago
a month ago
Title: 9x higher private-network latency in one environment vs another, same project, same region
I have two environments in the same project, both eu-west. Identical code and service setup. Production is roughly 5x slower end to end, and I've traced it to private-network round-trip latency.
ICMP to the MySQL service, from the app container in each environment. No TCP, no application code involved:
PROD app -> mysql-xxxx.railway.internal
min/avg/max = 4.512/4.643/4.934 ms (20 packets, 0% loss)
STAGING app -> mysql-xxxx.railway.internal
min/avg/max = 0.373/0.497/1.473 ms (20 packets, 0% loss)Both ttl=62, so the same number of hops. Prod is 9.3x slower with very little jitter, which reads as a fixed path difference rather than congestion.
In both environments the app and the database resolve to the same /64:
PROD app fd12:xxxx:xxxx:1:b000::... mysql fd12:xxxx:xxxx:1:4000::...
STAGING app fd12:xxxx:xxxx:1:b000::... mysql fd12:xxxx:xxxx:1:b000::...The containers themselves are equivalent, so this isn't a resource difference:
PROD STAGING
CPU benchmark (3M iter) 20ms 19ms
TCP connect to loopback 0.023ms 0.021ms
cpu.max 300000 100000
nr_throttled 0
memory 181MB / 3GB usedThe same gap shows up against a second, unrelated service:
PROD STAGING
MySQL, warm connection 9.53 ms/query 0.58 ms/query
Redis GET, warm 5.46 ms/get 0.28 ms/getMySQL benchmark() server-side CPU is 38.8ms vs 27.5ms, so the database instances perform comparably. The difference is entirely in getting packets there and back.
Impact: a typical page makes ~70 round trips across MySQL and Redis. At 4.6ms that's ~320ms of pure network latency per request, against ~35ms in staging. Server-timing on the same route, same query count (21 queries):
PROD total 555ms php 333ms db 222ms
STAGING total 51ms php 36ms db 15msIs there anything I can do to get the prod environment's services co-located the way staging's are? Would redeploying the database and Redis services trigger re-placement, or is placement fixed once a service is created?
One other thing I noticed in both environments: DNS resolution of *.railway.internal takes around 8-9ms. That seems high for an internal resolver and it's an extra cost on every new connection.