3 months ago
Summary
Our service behind the Railway edge (server: railway-hikari, x-railway-edge: atl1) returns 200 for a request, but the edge intermittently kills the HTTP/2 stream to the client with RST_STREAM(CANCEL) instead of delivering the response.
It only affects Android OkHttp clients over HTTP/2 — curl, Python (HTTP/1.1), and Node.js (HTTP/2) never reproduce it from the same networks. Since OkHttp refuses to retry POSTs on a reset stream (correctly, they're non-idempotent), every occurrence surfaces as a hard failure in our mobile app: users cannot sign up while our backend logs nothing but 200s.
Timestamp-matched evidence
Domain: core-api-prd.xxx (custom domain → Railway service running nginx in front of Node).
Example occurrence, 2026-07-16 04:54:29 UTC:
-
Origin (nginx access log, from Railway logs):
2026-07-16T04:54:26.443Z "POST /v1/users/check-email HTTP/1.1" 200 16 "-" "okhttp/4.12.0" <- delivered OK 2026-07-16T04:54:29.392Z "POST /v1/users/check-email HTTP/1.1" 200 16 "-" "okhttp/4.12.0" <- client got RST_STREAM(CANCEL) -
Client (Android, OkHttp 4.12): the second request fails ~325 ms after send with
okhttp3.internal.http2.StreamResetException: stream was reset: CANCEL(surfaces in React Native/axios as a bodyless "Network Error",
net::ERR_FAILED)
The same client had a successful 200 over the same connection 5 seconds earlier, so this is not an idle/stale-socket problem. Failed attempts consistently die ~300–335 ms in — a full round trip — matching "origin answered, edge canceled the delivery."
We also see this in production against real users: 17 Sentry events on 2026-07-16 between 03:27–03:47 UTC (POST /v1/users/check-email and /v1/auth/otp/verify), all Android, all matching nginx 200 log lines. Two different physical devices on two unrelated networks (one on residential IPv6, one on IPv4 Wi-Fi) reproduce it.
What we ruled out (all from the same egress IP as the failing device)
- Origin behavior — nginx logs 200 and delivers the response to the edge; no 5xx anywhere.
- nginx keep-alive recycling — a 1,326-request burst over a single connection (crossing the
keepalive_requests 1000boundary) completed with 0 failures. - Idle-connection boundaries — a 15-min steady HTTP/2 session (request every 10 s) and 40 fresh-connection POSTs at 90 s intervals: 0 failures.
- Request contents — replaying the failing app's exact header set (incl. long
baggage/sentry-traceheaders) over Node HTTP/2, 15/15 success. - Client network — the failing device resolves and reaches the edge normally (same LAN as passing curl probes; no proxy/VPN/private DNS).
The only variable we could not replicate outside a real device is OkHttp's TLS/HTTP-2 client fingerprint — every failure is OkHttp-h2; nothing else ever fails.
Separate, related measurement
While investigating we also measured that the edge closes idle client keep-alive connections silently at ~60 s (reuse works at 55 s idle, dead at 65 s — raw socket probe). Well-behaved clients that health-check pooled sockets likely survive this, so we're reporting it as an observation, not a confirmed failure cause. We've worked around both behaviors client-side (short OkHttp pool keep-alive + pinning HTTP/1.1, which eliminates the RST_STREAM issue entirely), but the h2 cancellation looks like an edge bug you'd want to know about.
Questions
- Is hikari known to reset h2 streams for OkHttp-fingerprinted clients (bot/fingerprint heuristics, h2-specific relay path)?
- Is there any per-service configuration to influence edge h2 behavior or idle timeouts?
- Can you trace the edge side of
POST xxx/v1/users/check-emailat 2026-07-16 04:54:29 UTC (edge atl1) to see why the downstream stream was canceled after the origin 200?
10 Replies
3 months ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 3 months ago
3 months ago
Im thinking about a regression of a previously confirmed Railway edge issue, rather than a new or purely OkHttp-side problem. And i would not currently attribute this to bot detection or TLS-fingerprint heuristics.
The nginx 200 16 entry is useful evidence, with one qualification: it proves that nginx produced the 200 response and wrote the response body toward Railway's proxy. It does not prove that Hikari successfully delivered the corresponding HTTP/2 HEADERS and DATA frames to the Android client. The suspected failure boundary remains between Railway's edge and OkHttp.
To make the escalation traceable, log Railway's request identifiers in nginx:
log_format railway_trace
'$time_iso8601 "$request" status=$status '
'body_bytes_sent=$body_bytes_sent '
'request_id=$http_x_railway_request_id '
'edge=$http_x_railway_edge '
'request_time=$request_time '
'upstream_response_time=$upstream_response_time '
'request_completion=$request_completion '
'user_agent="$http_user_agent"';
access_log /dev/stdout railway_trace;For each occurrence, provide Railway with:
X-Railway-Request-Id;- exact UTC timestamp;
- edge region;
- deployment ID;
- HTTP path and method;
- source IP, if it can be shared privately;
- whether the same HTTP/2 connection handled a preceding successful stream;
- the successful and failed HTTP/2 stream IDs.
OkHttp's HTTP/2 frame logging would provide the strongest client-side evidence. It distinguishes inbound frames with << and outbound frames with >>. On an affected device it should show whether the client receives something similar to:
<< ... RST_STREAM CANCELOn Android, OkHttp documents enabling this with:
adb shell setprop log.tag.okhttp.Http2 DEBUG
adb logcat '*:E' 'okhttp.Http2:D'That would establish both the stream ID and that the reset arrived from the remote peer, rather than being produced by an application-level Call.cancel().
It would also be useful to run the same Android build against both:
- the custom domain;
- the Railway-provided
*.up.railway.appdomain.
If both exhibit the same inbound RST_STREAM(CANCEL), an additional DNS/CDN layer can be ruled out.
ve-jo
Im thinking about a regression of a previously confirmed Railway edge issue, rather than a new or purely OkHttp-side problem. And i would not currently attribute this to bot detection or TLS-fingerprint heuristics. The nginx `200 16` entry is useful evidence, with one qualification: it proves that nginx produced the 200 response and wrote the response body toward Railway's proxy. It does not prove that Hikari successfully delivered the corresponding HTTP/2 HEADERS and DATA frames to the Android client. The suspected failure boundary remains between Railway's edge and OkHttp. To make the escalation traceable, log Railway's request identifiers in nginx: ```nginx log_format railway_trace '$time_iso8601 "$request" status=$status ' 'body_bytes_sent=$body_bytes_sent ' 'request_id=$http_x_railway_request_id ' 'edge=$http_x_railway_edge ' 'request_time=$request_time ' 'upstream_response_time=$upstream_response_time ' 'request_completion=$request_completion ' 'user_agent="$http_user_agent"'; access_log /dev/stdout railway_trace; ``` For each occurrence, provide Railway with: * `X-Railway-Request-Id`; * exact UTC timestamp; * edge region; * deployment ID; * HTTP path and method; * source IP, if it can be shared privately; * whether the same HTTP/2 connection handled a preceding successful stream; * the successful and failed HTTP/2 stream IDs. OkHttp's HTTP/2 frame logging would provide the strongest client-side evidence. It distinguishes inbound frames with `<<` and outbound frames with `>>`. On an affected device it should show whether the client receives something similar to: ```text << ... RST_STREAM CANCEL ``` On Android, OkHttp documents enabling this with: ```text adb shell setprop log.tag.okhttp.Http2 DEBUG adb logcat '*:E' 'okhttp.Http2:D' ``` That would establish both the stream ID and that the reset arrived from the remote peer, rather than being produced by an application-level `Call.cancel()`. It would also be useful to run the same Android build against both: 1. the custom domain; 2. the Railway-provided `*.up.railway.app` domain. If both exhibit the same inbound `RST_STREAM(CANCEL)`, an additional DNS/CDN layer can be ruled out.
3 months ago
Thanks for the suggestions — we ran all of them today. Results:
1. Alternative domain test: reproduced identically on the Railway-provided domain.
Same failure rate (~1 in 3 POSTs) on xxx.up.railway.app as on our custom domain. Custom domain / DNS / CDN factors are ruled out — this is internal to the edge.
2. Client-side capture: we instrumented OkHttp with an EventListener and caught the full failure narrative on-device. A representative failure (device clock, UTC-3):
10:04:01.686 responseHeadersStart <- same h2 connection serves a 200 normally
10:04:06.695 connectionAcquired h2 <- SAME connection, 5s later
10:04:06.696 requestHeadersEnd <- request headers fully sent
10:04:06.698 requestBodyEnd 41B <- body fully sent
10:04:06.945 responseHeadersStart <- the response STARTED arriving
10:04:06.9xx StreamResetException: stream was reset: CANCELThe response was already being delivered (headers arriving ~250 ms after send, matching our origin's processing time, which logged a 200) when the stream was reset mid-delivery. So the edge accepted the request, relayed it, received the origin's response, began forwarding it, and then canceled the stream.
3. Railway request IDs: captured for the successful requests bracketing each failure (canceled streams never complete headers, so they carry no id). Examples from today, edge atl1: olhv_e_uTZCwW2WP9o6EoQ, RDqwtUkjSQO9-yUm2prcFg, 0EplyObQSFixv0QyljLL4A — the failures occurred within seconds of these, from the same client connection. We're also adding x-railway-request-id + upstream timings to our origin's access log so future failures can be correlated server-side.
4. Control experiment: pinning OkHttp to HTTP/1.1 makes the failures unreproducible; re-enabling HTTP/2 brings them back within minutes at ~33% of POSTs. Same app, same device, same network, same origin — the only variable is the client↔edge protocol.
3 months ago
Follow-up with the exact request id of a canceled delivery, now that our origin logs x-railway-request-id:
rid=CLL4rFSIQOWSdceH6WHkDg — edge atl1, 2026-07-16 13:35:48 UTC.
Origin access log for that request:
[16/Jul/2026:13:35:48 +0000] "POST /v1/users/check-email HTTP/1.1" 200 16 "okhttp/4.12.0" rid=CLL4rFSIQOWSdceH6WHkDg rt=0.011 urt=0.011Our origin served it in 11 ms with a 200. The client (Android OkHttp 4.12, HTTP/2) had fully sent the request, saw response headers start arriving, and then received RST_STREAM(CANCEL) mid-delivery — captured by an OkHttp EventListener on the device at the same timestamp.
For contrast, the three requests immediately before it from the same device completed normally: rid=Ed4KDkYKQsy4cyS99o6EoQ (13:35:21), rid=rrTPUBj5TjCQGmSvxtoGcA (13:35:27), rid=S7KZMw-DThSl8kfx6WHkDg (13:35:32).
One more observation from the new logs that may help: each of these requests arrived at our origin on a fresh edge→origin connection ($connection/$connection_requests = 29/1, 31/1, 33/1, 35/1) — the edge doesn't appear to be reusing upstream connections at all here.
This should be directly traceable on your side now. Happy to provide more instances on demand — we can reproduce at ~1 in 3 POSTs.
3 months ago
One useful detail about the OkHttp timeline: in OkHttp 4.12, "responseHeadersStart" is not emitted merely because the client is ready to read.
Since OkHttp 4.3, it is emitted only after response-header bytes have actually started arriving.
Therefore, your trace does establish that the response had begun crossing the client-facing HTTP/2 connection before the stream failed.
There is only one remaining protocol-level caveat: an OkHttp "StreamResetException(CANCEL)" and an "EventListener" timeline do not independently identify which endpoint transmitted the "RST_STREAM". A frame log showing:
<< ... RST_STREAM CANCEL
would establish the direction conclusively.
However, I do not think that should block escalation anymore.
At this point, the remaining investigation requires Railway to inspect Hikari's internal trace for that request ID.
2 months ago
We have now independently reproduced what appears to be the same issue on a completely separate Railway service/application.
Our stack:
React Native / Expo Android app
Physical Google Pixel 10a, Android 17
OkHttp 4.9.2
Fastify / Node backend hosted on Railway
PostgreSQL mutation behind the affected request
We reproduced it twice on the same POST endpoint.
In both occurrences:
the request reached Fastify;
the database mutation was successfully committed;
Fastify completed with HTTP 201 in 8–18 ms;
Railway Network Logs also showed HTTP 201 with no upstream error;
approximately 150–250 ms later, the Android transport failed with:
stream was reset: CANCEL
React Native consequently surfaced this as TypeError: Network request failed, even though the operation had completed successfully.
Two concrete reproductions from the device:
mobile_msqqzc82_72e7pne4
Fastify: 201 in 17.74 ms
persisted: yes
client transport: stream was reset: CANCEL
mobile_msqr42di_a5ru109a
Fastify: 201 in 8.69 ms
persisted: yes
client transport: stream was reset: CANCEL
We also verified that our application is not cancelling these requests: there is no AbortController, mutation timeout or application-level cancellation on this path.
A PUT request in the same test session completed normally, while both POST attempts above failed this way.
Railway reports HTTP/1.1 upstream and HTTP/2 downstream for the affected traffic.
We initially couldn't see the underlying error because React Native collapses the native transport failure into Network request failed. We added Android-side instrumentation around the native XHR path and recovered the original OkHttp message, which is consistently stream was reset: CANCEL.
We have also added end-to-end request correlation to our Fastify service and can provide further Railway request IDs/timestamps after deploying that instrumentation.
This appears to independently confirm the behaviour reported in this thread, including that the origin successfully completes the response before the Android HTTP/2 client observes the stream reset.
If useful, we can provide additional reproductions, exact UTC timestamps, Railway Network Log details and Android transport logs.
a month ago
Same issue here, on a different edge region.
Setup
- NestJS/Express on Railway
- Edge: mia1 (the other report is atl1)
- Custom domain
- Client: React Native 0.81.5 on Android, using OkHttp
Observed behavior
The origin completes the request successfully with:
HTTP 200
Content-Length: 58
However, on the Android client we receive: "StreamResetException: stream was reset: CANCEL". The reset occurs after the response headers are received but before any response body bytes reach the application.
From Axios/Sentry, this surfaces as:
AxiosError: Network Error
ERR_NETWORK
hasRequest: true
hasResponse: false
xhrStatus: 200
nativeMessage: stream was reset: CANCEL
Production data. Between Aug 19 and Aug 25, we observed:
- 21 failures
- 7 users
- 7 different device models
- Android 12 through 16
- 9 different ISPs
- Both Wi-Fi and cellular
a month ago
Same issue.
Setup:
- Django on Railway
- Edge: sin1 (other than the previous two)
- Custom domain
- Client: React Native 0.81.5 on Android, using OkHttp
Observed behavior:
The backend successfully completes the request and responds with HTTP 200.
However, on the Android client, we receive: "Network Request failed"
We also tried HTTP 1.1, and the issue was not reproducible on the physical testing phone. When we switched back to HTTP 2, it started failing again.
It happens almost always for all users.
Although GET requests are working fine, only POST APIs are failing.
- Different ISPs from different countries/continent/region.. so it's not restricted to just one area.
a month ago
Fifth edge region, and by some margin the worst rate reported on this thread so far. lhr1, AdonisJS 6 on Node, React Native 0.81 / OkHttp 4.12.0 on Android.
Same signature as filipe's original report: origin answers 200, client gets StreamResetException: stream was reset: CANCEL, only POSTs carrying a request body, and pinning OkHttp to HTTP/1.1 makes it vanish entirely.
What's new here
1:1 correlation against Railway's own HTTP log. The device clock ran 2.01s behind Railway's. Applying that offset, every failed client request lines up with an edge row:
CLIENT (device clock) RAILWAY HTTP LOG (+2.01s)
10:53:50.904 CANCEL 10:53:52.919 200 65ms
10:53:51.700 CANCEL 10:53:53.707 200 58ms
10:53:52.151 CANCEL 10:53:54.160 200 58ms
... 30 attempts, 30 rows, 1:1, no gaps ...
10:54:06.144 CANCEL 10:54:08.155 200 58ms30 client requests produced 30 edge rows and 29 lost responses. A 97% failure rate in that run, against the ~33% in the original report. Every row: httpStatus 200, upstreamErrors empty, totalDuration 56-74ms.
Timing does not discriminate. The one request in that run that did reach the device logged totalDuration 56ms, sitting dead centre of the failure band. So totalDuration is useless for spotting an affected request from the log side. Only txBytes could separate "never written downstream" from "written and lost", and that field is not exposed through the Railway MCP server or the Railway agent.
Not sticky per connection. That single success landed in the middle of the 29 failures, on the same pooled connection that failed either side of it. Whatever fails, fails per-response rather than poisoning the connection.
Other HTTP/2 clients on the same line do not reproduce it. From the same network and the same endpoint: Node http2 ran 40 POSTs with bodies on one connection (40/40 OK), and 10 rounds of POST plus two concurrent GETs plus POST (40 streams, 0 failures). curl --http2 ran 10 POSTs with bodies on a reused connection, num_connects 0 on requests 2 through 10, 10/10 OK. So this is not the network path and not HTTP/2 generally. It is specifically OkHttp's HTTP/2 against hikari.
HTTP/1.1 on the same endpoint is clean. 10 POSTs with JSON bodies, 2s apart, on a single HTTP/1.1 connection: 10/10 succeeded. After pinning Protocol.HTTP_1_1 in the app: 15 client requests to 15 edge rows, 1:1, zero lost responses, no recurrence.
No CDN in the path. This reproduces on a plain *.up.railway.app domain with nothing in front of it, which removes the custom-domain variable present in the earlier reports.
Also ruled out
The edge correctly strips Connection, Keep-Alive and Transfer-Encoding on both the h1 and h2 downstream paths, so the "remove hop-by-hop headers from your responses" advice that circulates for this class of problem does not apply. The documented 60s HTTP/1.1 idle close is not it either, since failures happen seconds after a successful request and h2 is exempt from that rule. And it is not a stale pooled connection, because the connection demonstrably carried the request all the way to the origin.
Environment
Edge: x-railway-edge lhr1, consistent across 8 of 8 samples
Origin zone: railway/europe-west4-drams3a
Service: notepad-api
Deployment: d89d8483-6951-40a6-9210-d026c2fb35c0
Domain: notepad-backend-production-ec4f.up.railway.appOn the frame logs
ve-jo asked for OkHttp frame logs to confirm the reset direction and I don't believe anyone has supplied them yet. To be straight about the limits of what is above: I cannot confirm the direction from the JS layer either. stream was reset: CANCEL is produced both when the peer sends a stream reset with the CANCEL code, and when OkHttp's own reader loop reaches a clean end-of-connection, which sets streamErrorCode to CANCEL and then closes every stream still open on that connection. mlfraga's observation that the reset arrives after response headers but before any body bytes points at a genuine stream-level reset rather than a connection teardown, but the frames are what would settle it.
Happy to flip a build back to h2 with frame logging enabled and capture that, and to supply the client source IP plus exact timestamps for log lookup, if it helps move this forward.
24 days ago
Hey all, we have rolled out a fix for this. It turned out to be an issue with an upstream dependency.
Please let us know if you continue to see issues.
Status changed to Awaiting User Response Railway • 24 days ago
17 days ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • 17 days ago
3 days ago
Hey! I did some tests, and it seems to be fixed. Thanks, guys!
Attachments
Status changed to Open Railway • 3 days ago

