a month ago
Summary:
An external webhook sender (TradingView) is intermittently failing to deliver POST requests to my Railway service's public domain. The sender reports the exact error: "Webhook delivery failed — couldn't find this domain." This is a DNS resolution failure — the sender cannot resolve my *.up.railway.app hostname to an IP, so the requests never leave the sender and never reach my service. Some requests to the identical URL succeed while others fail, including alerts fired in the same second.
When it started:
This began today. The exact same setup, same URL, and same request volume/burst pattern have worked reliably every prior day. Nothing changed on my side (no deploys, no DNS/domain changes, no code changes to the webhook) around when it started.
What I have already checked and ruled out (please don't re-request these — they're confirmed):
My application is not erroring. Railway Metrics show Request Error Rate at a flat 0.0% over the affected period. No 4xx/5xx in the Requests graph.
No resource pressure. CPU is near-idle (~0.05 of 0.8 vCPU). Memory is flat (~300 MB). Response Time p50–p99 is well within normal (sub-second). The service is not overloaded.
The failing requests never reach my service. They do not appear in either my Deploy Logs or Network Logs (HTTP or DNS tabs). Successful requests are logged normally; failed ones have no entry at all — consistent with the requests dying at DNS resolution before transit.
No deployments/restarts during the failures. The failures are not correlated with any redeploy or restart; the service shows "Online / Active" throughout.
The sender's platform is operational. TradingView's status page shows all systems (including Alerts) operational, so this is not a sender-side outage.
The URL is correct and identical across all alerts. Successful and failed deliveries use the exact same webhook URL. The failures are intermittent, not tied to a specific malformed URL.
The sender's own error message is explicitly a domain-resolution failure ("couldn't find this domain"), not a timeout, connection refused, or HTTP error.
Conclusion / likely cause:
This points to intermittent DNS resolution failures for my *.up.railway.app subdomain — i.e., the DNS serving up.railway.app is intermittently not returning a valid answer to the sender's resolver. Because it's at the name-resolution layer, requests never reach my container (hence 0% error rate and no logs), and it's intermittent (same-second alerts split between success and failure).
My questions / requests:
Is there a known or ongoing DNS/resolution issue affecting *.up.railway.app subdomains, or my service's region specifically, today?
Can you check the DNS/edge health for my service's public domain from your side and confirm whether resolution is intermittently failing?
Is there anything on my end or yours that can stabilize resolution for the Railway-provided subdomain?
As a mitigation, I am planning to add a custom domain fronted by a third-party DNS provider (Cloudflare) and repoint the sender to it. Do you have any guidance or known issues with custom domains + webhook traffic that I should be aware of, and would this be expected to bypass the *.up.railway.app resolution path entirely?
This is a production trading system where dropped webhooks have direct financial impact, so a timely response would be greatly appreciated. Happy to provide screenshots of the metrics, logs, and the sender's error messages.
1 Replies
a month ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 1 month ago
a month ago
Here is the step-by-step solution I would follow:
1. Confirm the DNS problem
Replace YOUR-RAILWAY-DOMAIN with your actual *.up.railway.app hostname:
dig YOUR-RAILWAY-DOMAIN @1.1.1.1
dig YOUR-RAILWAY-DOMAIN @8.8.8.8Run each several times. Check whether you get inconsistent results, SERVFAIL, NXDOMAIN, or timeouts.
2. Test from another network
Repeat the same commands from another machine, VPS, or mobile hotspot. This rules out a local DNS problem.
3. Add a custom domain
In Railway:
Service → Settings → Networking → Custom Domain
Add something like:
webhook.yourdomain.com
Railway will provide the required DNS records. Add those records to your DNS provider/Cloudflare exactly as Railway shows them.
4. Verify the domain
Wait until Railway shows the custom domain as verified and HTTPS is active.
Test:
dig webhook.yourdomain.com
curl -v https://webhook.yourdomain.com/your-webhook-path5. Test the actual webhook
Send a test POST:
curl -X POST https://webhook.yourdomain.com/your-webhook-path \
-H "Content-Type: application/json" \
-d '{"test":true}'Confirm the request appears in Railway Network Logs.
6. Change TradingView
Once the custom domain works consistently, change the TradingView webhook URL from:
https://YOUR-RAILWAY-DOMAIN/...
to:
https://webhook.yourdomain.com/...
Keep the old Railway URL active temporarily as a fallback.
7. If the custom domain works
Continue using the custom domain for production. Send Railway the dig results and timestamps of the failed requests so they can investigate the *.up.railway.app DNS/edge issue.
8. If the custom domain also fails
Then the problem is probably not limited to Railway's up.railway.app DNS. Provide Railway with the DNS tests, curl -v output, exact failure timestamps, and TradingView delivery errors so they can investigate the edge/network layer.
If the issue still happens after switching to the custom domain, don't keep changing DNS settings blindly. Collect the exact failure timestamps, dig results, curl -v output, and TradingView error screenshots, then open a Railway support ticket and ask them to investigate the DNS/edge layer. If only TradingView fails while your own curl tests consistently work, also provide the same evidence to TradingView support so they can check their DNS resolver/egress path.