a month ago
Hi Railway team,
I'm seeing an intermittent issue where requests to my backend service are processed successfully, but the response doesn't reliably reach the client. I'd like your help ruling out anything on Railway's infrastructure side.
Setup:
- Plan: Hobby
- Service: Node.js/Express backend (Mongoose/MongoDB)
- Client: React Native mobile app, primary user base in Nigeria (mostly cellular networks)
Symptom:
- Server logs confirm the request is received, processed, and a 200 response is sent (e.g. writes complete successfully, confirmed via application logs).
- The mobile client sometimes never receives that response. React Native reports this as a generic "TypeError: Network request failed" — i.e., the connection appears to drop before the response arrives, not a timeout (our client timeout is 30s; server processing consistently takes under 3s when tested directly).
- This happens across multiple endpoints (POST and PATCH requests), not one specific route.
What I've already ruled out:
- Serverless/sleep mode is confirmed off for this service, so this isn't a cold-start/wake delay.
- Server response time is consistently ~3s when tested directly (via Bruno), well under any timeout window.
- No usage limit or hard cap appears to be in effect that would explain a mid-cycle shutdown.
What I'd like your help with:
- Are there known idle-connection or keep-alive timeout mismatches between Railway's edge proxy and backend services on the Hobby plan that could cause a response to be cut off mid-transfer?
- Is there anything region-specific I should consider — my service is currently deployed in [REGION], and my users are in Nigeria/West Africa, so there's meaningful distance/hops involved. Would a different available region reduce intermittent drops, or is this unlikely to be proxy-related at all?
- Are there any logs or metrics on your side (connection resets, proxy-level errors) for my service around the time these drops occur that could help confirm whether this is happening on Railway's edge vs. further out on the network path?
Happy to share more detail — just let me know what would be useful.
Thanks for your help
5 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 26 days ago
a month ago
To answer your questions:
- Railway allows a 60s idle connection
- You mention
[REGION]but not what your region actually is, considering your user location Europe will be the best region though, worth a try but I wouldn't hedge the success of a fix on a region change. A slightly slow network (which is what a region change would help solve) shouldn't be causing non deterministic errors in your app - All network related logs can be found under Deployment -> Network Logs -> HTTP / DNS / Net Flow
What I would do in your shoes is I'd log the Railway request ID and Railway Edge and compare those with network logs to see if I can't find any leads
And in the mean time I'd add client retry (assuming your POST and PATCH requests allow for this), this should patch the issue until you have a more permanent fix
Does every single failure match a 200 response from the server?
dev
To answer your questions: 1. Railway allows a 60s idle connection 2. You mention `[REGION]` but not what your region actually is, considering your user location Europe will be the best region though, worth a try but I wouldn't hedge the success of a fix on a region change. A slightly slow network (which is what a region change would help solve) shouldn't be causing non deterministic errors in your app 3. All network related logs can be found under Deployment -> Network Logs -> HTTP / DNS / Net Flow What I would do in your shoes is I'd log the Railway request ID and Railway Edge and compare those with network logs to see if I can't find any leads And in the mean time I'd add client retry (assuming your POST and PATCH requests allow for this), this should patch the issue until you have a more permanent fix Does every single failure match a 200 response from the server?
25 days ago
Hello, thank you for the detailed reply.
To answer your question: it's intermittent, not consistent. My region was first set to US EAST, then I changed it to EU WEST, and I'm still having this issue. Most of the time, POST and PATCH requests succeed cleanly on both the client and server—no issue at all. But every so often, within a span of a few minutes of otherwise normal usage, the client will suddenly report a failure even though the server logs show a 200 response was sent successfully for that same request. It's not tied to a specific endpoint; I've seen it on more than one POST/PATCH route—and I haven't yet found a pattern in when it happens (not obviously tied to time of day, specific requests, or load).
So the core mystery for me right now is exactly that: the server is confirming success (200, write completed), but the client is treating it as a failure, and I don't yet know why the client is doing that on those specific occasions when everything looks fine moments before and after. I am using Tanstack Query in my React Native app.
25 days ago
Based on your symptom set, the likely issue is not that your server fails to process the request. It’s that the response is lost after processing, either at the Railway edge, on the mobile carrier path, or during a Node/Railway keep-alive mismatch. Because the server already committed the write, the real fix is to make retries safe and to prove where the response dies.
You can try idempotent writes and client retry ... I think for every POST/PATCH that can be retried, have the React Native client send an idempotency key. The server stores the result for that key and replays the same response if the client retries. This solves the user-facing problem: if the phone never received the 200, it can safely retry and get the original success response instead of creating a duplicate.
The client should retry on network failure with backoff and jitter. Your 30s timeout is fine; the issue is not timeout, it’s a dropped connection. Retrying on TypeError: Network request failed is the correct behavior, but only if the server is idempotent.
Or try aligning the Node keep-alive with Railway’s edge
Plus i'd advice moving the Railway service to EU West / Amsterdam if it isn’t already there. That is the closest Railway region to Nigeria/West Africa. US East is second. This reduces latency and packet loss probability, but it will not reliably fix a response that is lost after your server returned 200. The idempotent retry is what actually fixes the user impact.
If you share the repo or invite me, I’ll pull it down and help fix this directly. I can also review your current Express setup and React Native API layer first and tell you exactly which endpoints need idempotency and which retry logic is safe.
clemon1
Hello, thank you for the detailed reply. To answer your question: it's intermittent, not consistent. My region was first set to US EAST, then I changed it to EU WEST, and I'm still having this issue. Most of the time, POST and PATCH requests succeed cleanly on both the client and server—no issue at all. But every so often, within a span of a few minutes of otherwise normal usage, the client will suddenly report a failure even though the server logs show a 200 response was sent successfully for that same request. It's not tied to a specific endpoint; I've seen it on more than one POST/PATCH route—and I haven't yet found a pattern in when it happens (not obviously tied to time of day, specific requests, or load). So the core mystery for me right now is exactly that: the server is confirming success (200, write completed), but the client is treating it as a failure, and I don't yet know why the client is doing that on those specific occasions when everything looks fine moments before and after. I am using Tanstack Query in my React Native app.
25 days ago
interesting, thanks for the detail
Do you happen to know the size of responses (or maybe the size of failed responses specifically) and the time to failure on the client?
25 days ago
Hey all, we have rolled out a fix for this. It turned out to be an issue with an upstream dependency.
Please let us know if you continue to see issues.
Status changed to Awaiting User Response Railway • 25 days ago
17 days ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • 17 days ago
