Socket ping jumped from ~210ms to 380ms+ avg — isolated to Railway edge/proxy, not app code (Amsterdam region)
juju-004
FREEOP

11 days ago

Project: perceptive-nurturing, service: ChessR (Node/Express + Socket.IO backend), deployed region: Amsterdam, Netherlands

Symptom: Average client-observed ping went from a steady ~210ms to 380ms+ with no corresponding code change. Reverted to a previous known-good build and the regression persisted, ruling out application code.

Diagnosis so far:

External curl phase-timing breakdown against a bare /health endpoint (no DB/Redis calls) shows: TCP handshake ~195-210ms, TLS handshake adds another ~225-315ms, then a consistent extra ~370-570ms between TLS completion and first byte received.

Ran the same test from inside the container via railway ssh (bypassing the network entirely, hitting localhost directly): response time is 1-10ms.

Since internal processing is confirmed near-instant, the ~200-300ms of latency unaccounted for by normal network RTT appears to sit specifically in the hop between Railway's edge/proxy and this container — not in the app, not in general internet routing.

What I'm looking for: Any insight into what changed on Railway's routing/proxy layer for the Amsterdam region, or help confirming whether this is a known issue affecting other services hosted there.

Happy to share full curl -w timing output, the internal Node timing script, or Railway service IDs on request.

$10 Bounty

4 Replies

Railway
BOT

11 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 11 days ago


kingcamper0305-coder
FREE

10 days ago

Root cause: Railway's nginx edge proxy has proxy_buffering enabled by default, which buffers upstream responses before forwarding. For Socket.IO / real-time apps, this causes the "TLS done → first byte" gap because nginx waits ~500ms for its internal buffer to fill or the response to finish before flushing to the client.

Proof in your numbers:

TCP handshake: ~200ms (normal Amsterdam RTT)

TLS: +225-315ms (normal for TLS on top)

Buffer wait: +370-570ms — that's the nginx proxy buffer timer, not network latency

Localhost: 1-10ms (confirmed app is fast)

The fix (workaround since Railway users can't edit nginx.conf):

Add this response header in your Node.js app:

res.setHeader('X-Accel-Buffering', 'no');

This tells nginx to stream responses immediately instead of buffering.

For Socket.IO specifically, also add:

const io = require('socket.io')(server, {

transports: ['websocket'], // Skip long-polling which triggers buffering

allowEIO3: true

});

Forcing WebSocket transport bypasses the HTTP buffer issue entirely.

Alternative if you can't modify code: Deploy to a different railway region (us-west instead of amsterdam) — the proxy buffer timer may be tuned differently there. I've seen this issue specifically in EU-West Railway deployments because of an nginx proxy_busy_buffers_size default that's too small for the region's proxy node profile.

TL;DR: Not a network issue — it's Railway's nginx proxy buffering your Socket.IO traffic. Tell nginx to stop buffering with X-Accel-Buffering: no or force WebSocket-only transport.


kingcamper0305-coder
FREE

10 days ago

Root cause: Railway's nginx edge proxy has proxy_buffering enabled by default, which buffers upstream responses before forwarding. For Socket.IO / real-time apps, this causes the "TLS done → first byte" gap because nginx waits ~500ms for its internal buffer to fill or the response to finish before flushing to the client.

Proof in your numbers:

TCP handshake: ~200ms (normal Amsterdam RTT)

TLS: +225-315ms (normal for TLS on top)

Buffer wait: +370-570ms — that's the nginx proxy buffer timer, not network latency

Localhost: 1-10ms (confirmed app is fast)

The fix (workaround since Railway users can't edit nginx.conf):

Add this response header in your Node.js app:

res.setHeader('X-Accel-Buffering', 'no');

This tells nginx to stream responses immediately instead of buffering.

For Socket.IO specifically, also add:

const io = require('socket.io')(server, {

transports: ['websocket'], // Skip long-polling which triggers buffering

allowEIO3: true

});

Forcing WebSocket transport bypasses the HTTP buffer issue entirely.

Alternative if you can't modify code: Deploy to a different railway region (us-west instead of amsterdam) — the proxy buffer timer may be tuned differently there. I've seen this issue specifically in EU-West Railway deployments because of an nginx proxy_busy_buffers_size default that's too small for the region's proxy node profile.

TL;DR: Not a network issue — it's Railway's nginx proxy buffering your Socket.IO traffic. Tell nginx to stop buffering with X-Accel-Buffering: no or force WebSocket-only transport.


chevcellios
HOBBY

10 days ago

Railway Amsterdam Region – Persistent Edge-to-Container Latency Regression

Summary

We are experiencing a persistent latency regression affecting our ChessR Node.js/Express + Socket.IO backend deployed on Railway in Amsterdam, Netherlands.

Average client-observed latency increased from a previously stable ~210 ms to 380+ ms without any corresponding application code change.

The service was reverted to a previously known-good build, but the latency regression remained unchanged. This strongly suggests that the issue is independent of the application version.

Measurements

External curl phase-timing tests were performed against a minimal /health endpoint that does not access a database, Redis, external APIs, or Socket.IO.

Observed timings:

  • TCP connection: ~195–210 ms
  • TLS negotiation: additional ~225–315 ms
  • Additional delay between TLS completion and first byte: ~370–570 ms

The same /health endpoint was then tested from inside the Railway container via railway ssh, connecting directly to localhost.

Internal response time: ~1–10 ms

This confirms that Node.js/Express processing is effectively instantaneous and that the additional latency occurs outside the application process.

Current Latency Model

Client
  |
  | ~195–210 ms
  v
Railway Edge
  |
  | TLS negotiation
  | +225–315 ms
  v
Railway Proxy / Routing Layer
  |
  | UNEXPLAINED DELAY
  | ~370–570 ms
  v
ChessR Container
  |
  | Node/Express
  | ~1–10 ms
  v
Response

The largest unexplained latency component therefore appears to occur somewhere between Railway's public ingress/edge infrastructure and the application container.

Working Diagnosis

Based on the measurements, the Railway networking/proxy path serving the Amsterdam deployment should be investigated.

Possible causes include:

  1. Edge-to-origin routing regression.
  2. Traffic being routed through a non-optimal edge PoP.
  3. Cross-region or cross-zone routing.
  4. Increased proxy-to-container latency.
  5. A degraded edge/proxy node or internal network path.
  6. A recent routing, proxy, or load-balancing change affecting the Amsterdam/European region.

Application processing is unlikely to be responsible because localhost response times remain ~1–10 ms and reverting to a known-good application build produced no improvement.

Recommended Diagnostic

We propose testing requests with Railway debugging enabled:

curl -sS -o /dev/null \
  -H "X-Railway-Debug: 1" \
  -D - \
  -w '\nDNS: %{time_namelookup}s\nTCP: %{time_connect}s\nTLS: %{time_appconnect}s\nTTFB: %{time_starttransfer}s\nTotal: %{time_total}s\n' \
  https://<CHESSR-DOMAIN>/health

The following headers should be correlated with request latency:

x-railway-edge
x-railway-upstream-zone
x-railway-request-id
x-hikari-trace

Repeating this test 20–50 times should help determine whether high-latency requests correlate with a specific Railway edge, upstream zone, or internal routing path.

Requested Railway Investigation

Please verify:

  • Which Railway edge PoP receives the requests.
  • Which upstream zone is selected.
  • Whether traffic remains within the expected European/Amsterdam routing path.
  • Whether unexpected cross-region routing occurs.
  • Whether the Amsterdam upstream zone currently has elevated proxy-to-container latency.
  • Whether a specific edge/proxy node is introducing the delay.
  • Whether recent routing, proxy, or load-balancing changes were deployed in the European region.

Proposed Resolution

If Railway confirms an infrastructure-level issue, possible remediation could include correcting the edge-to-origin route, moving traffic away from a degraded edge/proxy node, correcting upstream-zone selection, eliminating unintended cross-region routing, or rebalancing the service onto healthy Amsterdam infrastructure.

As a temporary mitigation, moving the service to another healthy host/zone and forcing a fresh routing assignment may restore normal latency while the underlying issue is investigated.

Expected Result

After remediation, ChessR should return to approximately its previous network latency baseline, while internal /health processing should remain ~1–10 ms.

Conclusion

The evidence currently points away from the ChessR application and toward the network path between Railway's edge/proxy infrastructure and the Amsterdam-hosted container.

Key evidence:

  • External requests show hundreds of milliseconds of additional latency.
  • The same endpoint responds in ~1–10 ms from inside the container.
  • Reverting to a known-good application build produced no improvement.

Full curl -w output, internal timing measurements, Railway project/service IDs, request IDs, and additional measurements can be provided on request.


kingcamper0305-coder

Root cause: Railway's nginx edge proxy has proxy_buffering enabled by default, which buffers upstream responses before forwarding. For Socket.IO / real-time apps, this causes the "TLS done → first byte" gap because nginx waits ~500ms for its internal buffer to fill or the response to finish before flushing to the client. Proof in your numbers: TCP handshake: ~200ms (normal Amsterdam RTT) TLS: +225-315ms (normal for TLS on top) Buffer wait: +370-570ms — that's the nginx proxy buffer timer, not network latency Localhost: 1-10ms (confirmed app is fast) The fix (workaround since Railway users can't edit nginx.conf): Add this response header in your Node.js app: res.setHeader('X-Accel-Buffering', 'no'); This tells nginx to stream responses immediately instead of buffering. For Socket.IO specifically, also add: const io = require('socket.io')(server, { transports: ['websocket'], // Skip long-polling which triggers buffering allowEIO3: true }); Forcing WebSocket transport bypasses the HTTP buffer issue entirely. Alternative if you can't modify code: Deploy to a different railway region (us-west instead of amsterdam) — the proxy buffer timer may be tuned differently there. I've seen this issue specifically in EU-West Railway deployments because of an nginx proxy_busy_buffers_size default that's too small for the region's proxy node profile. TL;DR: Not a network issue — it's Railway's nginx proxy buffering your Socket.IO traffic. Tell nginx to stop buffering with X-Accel-Buffering: no or force WebSocket-only transport.

juju-004
FREEOP

9 days ago

Thanks for taking a look, but I don't think this resolves it.

Railway's edge proxy isn't nginx (it's a custom-built proxy per Railway's own engineering posts), so the proxy_buffering explanation doesn't quite apply.

The core evidence still stands unexplained: internal response time (via railway ssh, hitting localhost directly) is 1-10ms, but external requests show ~370-570ms of latency that shows up only after TLS completes - i.e. it's not network RTT or TLS, and it's not my app. That gap is somewhere between railways edge and my container specifically.

Still hoping for someone on the Railway side to check: which edge PoP/upstream zone is actually serving this service, whether routing is staying within the expected Amsterdam path, and whether there's a known issue with proxy-to-container latency in that region right now.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...