3 months ago
Hi Railway Team,
We are experiencing high latency on our production application (FlexyBank) hosted on Railway.
Since yesterday, navigating between pages has become noticeably slower. API requests eventually complete successfully, but each page transition takes much longer than normal.
Yesterday we noticed there was an incident affecting the US region, and the issue still seems to persist for us today.
Our observations:
Deployment is healthy.
No application errors or crashes.
Database is responding normally.
The slowdown affects page navigation and API response times intermittently.
Multiple users are experiencing the same issue from different locations.
Could you please confirm if there is still an ongoing infrastructure issue affecting the US region or our deployment?
If needed, we can provide additional logs or diagnostics.
Thank you for your support.
2 Replies
3 months ago
This thread has been opened as a public bounty so the community can help solve it. The thread and any further activity are now visible to everyone.
Status changed to Open Railway • 3 months ago
3 months ago
A few things I'd check before concluding it's a Railway infrastructure issue.
Since your deployment is healthy and the database appears to be responding normally, try narrowing down where the latency is occurring:
- Check your service metrics (CPU, memory, and network) during the slow periods.
- Review your application logs to see whether the requests are spending time inside your app or before they reach it.
- If you're using multiple Railway services (API, database, Redis, etc.), verify whether only one service is experiencing increased latency.
- Run a few
curlrequests directly against your API from another machine or region to compare response times with what users are seeing in the browser. - If you have tracing or APM (OpenTelemetry, Sentry, Better Stack, etc.), see where the additional latency is being introduced.
Regarding yesterday's US region incident, I haven't seen confirmation that it's still ongoing. You may want to check whether the latency is reproducible across different regions or from an external monitoring service like UptimeRobot or Better Stack.
If you can share:
- your deployment region,
- a rough before/after response time (e.g. 200 ms → 3 s),
- whether you're on Hobby or Pro,
- and whether your frontend and backend are both hosted on Railway,
it'll be easier for us to help narrow down the cause.
3 months ago
Hi,
Thanks for your response.
We investigated further and this appears to be a general latency issue across most page navigations, not only one specific endpoint.
Details:
Plan:
- Pro
Frontend:
- Hosted on cPanel
Backend:
- Hosted on Railway
- Region: US
PostgreSQL:
- Railway PostgreSQL
- Region: US
Redis:
- Railway Redis
- Region: US
Context:
- Backend deployment is healthy.
- PostgreSQL is online.
- Redis is online.
- No crash loop.
- This started after the recent Railway US region incident.
- Before the incident, page navigation felt much faster.
- After the incident, most page transitions now show loaders for around 2–4 seconds, and some DB-backed admin endpoints take longer.
Direct backend measurement from my machine:
GET /health:
DNS: 0.002s
Connect: 0.295s
TLS: 0.482s
TTFB: 0.819s
Total: 0.820s
GET /health/db:
HTTP 200
dbMs: 440ms
Application logs show general slow requests:
- /api/auth/me → totalMs around 1050ms
- /api/public/tenant/resolve → totalMs around 1188ms
- /api/admin/notifications/unread-count → queryMs around 294ms, but totalMs around 2073ms
- /api/admin/dashboard → totalMs around 2963ms
- /api/admin/branches → totalMs around 3819ms
- /api/admin/branches/:id/financial-position → totalMs around 8324ms
We understand some endpoints can be optimized, and we have already started instrumenting/optimizing obvious slow queries. However, the issue is broader: even lightweight/common endpoints like auth/me, tenant resolve, and unread-count are slower than before.
Could you please check if there is still degraded latency in the US region, Railway internal networking, PostgreSQL connectivity, or edge-to-service routing?
Also, can you confirm whether our backend, PostgreSQL, and Redis are actually running in the same physical region/zone and not only the same broad US region?
One more observation:
From curl response headers we sometimes see Railway edge like ams1, while our services are in the US. Could this edge routing add extra latency for users outside the US, or is this expected?
Thanks.