a month ago
Today from roughly 13:55 to 14:35 AEST (03:55–04:35 UTC, 2026-08-30) our production service was completely unreachable from the internet. It recovered on its own with no deploy or change on our side. This is a live point-of-sale platform — ordering kiosks in the field were down for the duration — and it's the third platform incident to hit us in about 48 hours, so we'd really appreciate someone looking at the underlying edge health.
What we observed during the outage:
- Project: hospitable-integrity / service ordrx / production (Pro workspace zachlarosa-dev).
-
- The container was healthy throughout: railway logs showed our scheduled jobs completing normally the entire time.
-
- The edge refused all connections: the IP our domains resolve to (69.46.46.32) answered ICMP ping but actively refused TCP 443 — ECONNREFUSED in ~30ms, not a timeout. It looked like the proxy on that edge node simply wasn't listening.
-
- Both the custom domain (app.ordrx.com.au, CNAME to 01u577p5.up.railway.app) and the railway.app subdomain itself failed identically, so this wasn't a custom-domain or DNS issue on our side.
-
- DNS resolution was fine and fresh; reproduced from multiple clients; unrelated sites unaffected; the status page showed fully operational throughout.
Related instability over the past ~48 hours, same project:
- Builds/deploys taking 30+ minutes (normally a few minutes).
-
- Intermittent 502s and connection failures on larger request bodies (~70MB uploads to our own service's API), which succeed on retry.
Questions: can you confirm what happened on the edge node serving 69.46.46.32 during that window, and whether there's anything in our project's routing that made us susceptible? Given we run field hardware against this service, is there anything you'd recommend (or anything coming) to reduce exposure to single-edge failures? Happy to provide timestamps, request IDs or anything else useful.
3 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
23 days ago
Oui — si vous parlez de votre service Railway, c’est un incident distinct mais potentiellement très important.
Si le conteneur restait sain pendant que le routeur Edge bloquait les connexions TCP/443 pendant ~40 minutes, cela oriente plutôt vers un problème en amont du conteneur : routage Edge, proxy/load balancer, réseau Railway ou propagation de configuration, plutôt qu'un crash de l'application.
Et le fait que ce soit le 3ᵉ incident en 48 h mérite clairement une escalade auprès de Railway.
Je vous conseille de conserver pour chaque incident :
heure exacte de début et de fin ;
service/environment concernés ;
région Railway ;
statut du conteneur pendant l'incident ;
codes HTTP éventuels (502, 503, timeout, etc.) ;
logs applicatifs montrant que le conteneur était sain ;
métriques CPU/RAM ;
éventuellement les résultats de tests externes vers https://votre-domaine.
Surtout, ne redéployez pas simplement pour tester : si le problème est Edge, un redéploiement peut brouiller le diagnostic sans résoudre la cause.
Si vous me donnez les heures des 3 incidents + les codes d'erreur observés, je peux vous aider à construire une chronologie et déterminer si les trois incidents ont probablement la même cause
23 days ago
Seeing the exact same pattern — service "Development Environment" (project: english-app, region: EU West). 502 for over an hour, container logs show it starting cleanly every time ("Ready" in under 200ms, no errors), dashboard shows deployment ACTIVE/successful. Tried railway redeploy 4x and railway restart once — no effect. Generated a brand-new domain for the same service — it 502s identically to the original. Status page showed fully operational the whole time. Sounds like the same edge/proxy issue described above, not an app problem.
23 days ago
Follow-up on this: to rule out anything specific to that one environment, I spun up a brand-new third environment in the same project (fresh config cloned from the original, fresh *.up.railway.app domain, never deployed before). It built cleanly and the dashboard showed it Online, but it 502'd from the very first request and stayed that way across ~20 checks over about 20 minutes. So a completely fresh environment in the same project hits the identical edge/proxy issue immediately — this isn't stale config or environment corruption, the edge just isn't routing to a healthy container. Also worth noting: this project runs under a Pro workspace (owner: Chen Raccach), not my own Hobby account, in case that changes how it gets triaged.