Robokassa callbacks time out before reaching our Railway service
gyt0r
HOBBYOP

a month ago

Hello Railway Support,

We are investigating failed payment callbacks to a Railway-hosted production service and need help tracing inbound HTTPS connectivity at the Railway edge.

Our payment provider, Robokassa, reports that it sends POST callbacks from:

185.59.216.65

185.59.217.65

For payment invoice it attempted delivery four times:

2026-08-30 06:48:25 UTC

2026-08-30 06:49:25 UTC

2026-08-30 06:50:25 UTC

2026-08-30 06:51:26 UTC

Every attempt timed out after 30 seconds on Robokassa’s side with:

Operation time out

The issue has been ongoing since August 28, 2026. Earlier callbacks worked normally; for example, an older payment received the expected OK response.

What we verified:

  • The Railway domain is active and targets port 8080.
  • Public DNS resolves correctly.
  • Port 80 is available and redirects to HTTPS.
  • HTTPS on port 443 is available and routes to the application.
  • The callback endpoint accepts external POST requests; a controlled invalid POST reached the application and received an immediate HTTP 400 response.
  • There are no Edge Rules with Block or Challenge actions.
  • Railway HTTP logs contain no POST /payments/robokassa/result records during the timestamps above.
  • Nest application logs contain no request-handler entry during these attempts.

Therefore, the callbacks appear not to reach the application or Railway HTTP-level logging.

  • We explicitly created Railway Edge Rules that allow inbound requests from both Robokassa IPs (185.59.216.65 and 185.59.217.65). There are no Block or Challenge rules.
  • Despite these allow rules, Railway HTTP logs still contain no callback requests from those IPs during the affected period.

At Robokassa’s request, we also tested outbound connectivity from the Railway deployment to their two documented IPs:

ping -c 4 -W 5 185.59.216.65 → 100% packet loss

ping -c 4 -W 5 185.59.217.65 → 100% packet loss

nc -vz -w 10 185.59.216.65 443 → Operation timed out

nc -vz -w 10 185.59.217.65 443 → Operation timed out

We understand these hosts may intentionally block ICMP and inbound TCP/443, so this does not by itself prove a Railway issue. However, together with the callback timeouts and the absence of HTTP records, it suggests a possible routing or edge-connectivity issue between Robokassa and Railway.

Could you please investigate edge/proxy/network-flow logs for inbound connections from 185.59.216.65 and 185.59.217.65 during the listed UTC timestamps? In particular, we need to know whether the connections reached Railway’s edge, were dropped before HTTP logging, or were affected by a specific edge POP or routing issue.

Thank you.

$10 Bounty

2 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


gyt0r
HOBBYOP

a month ago

We performed a real callback capture with webhook.site. Robokassa successfully delivered the POST from source IP 185.59.216.65 in 1 ms. The request uses HTTP/1.1, application/x-www-form-urlencoded; charset=utf-8, .NET Framework/v4.0.30319, and does not include Expect: 100-continue. Therefore, the Expect: 100-continue hypothesis is ruled out.

The same provider source IP times out only when targeting our Railway domain, and no Railway HTTP log is created.


arthurveisseire
PRO

a month ago

Two things in the report are dead ends, and dropping them narrows this a lot.

Edge Rules act on HTTP requests that already arrived at the edge. A connection that never completes never becomes an HTTP request, so allow rules can neither help nor be tested this way. That's consistent with your own finding that no HTTP log exists for those timestamps: there was nothing at the HTTP layer to rule on. Same for the outbound ping/nc from your container to Robokassa's addresses, which measures the opposite direction and says nothing about their inbound path. You suspected as much already, and I'd cut both from the report so they don't muddy it.

What's left is TCP/TLS from 185.59.216.65 to the single IPv4 address your hostname resolves to. A 30-second "operation timed out" (not a reset, not a TLS alert) with no HTTP record on Railway's side means the connection isn't completing at all, which is below anything configurable in the dashboard.

The useful part: Railway hands out one A record per hostname from its own /24, but the whole /24 is a single shared front end, and every address in it serves every hostname.

$ whois 69.46.46.46 | grep -E 'NetRange|NetName|OrgName'
NetRange:       69.46.46.0 - 69.46.46.255
NetName:        RLWY-HIKARI-01
OrgName:        Railway

$ dig +short docs.railway.com
69.46.46.46
(other Railway hostnames land on .22, .69, .98, .110 ...)

$ for ip in 69.46.46.22 69.46.46.69 69.46.46.98 69.46.46.110; do
    curl -so /dev/null -w "$ip -> %{http_code}\n" --resolve docs.railway.com:443:$ip https://docs.railway.com/
  done
69.46.46.22 -> 200
69.46.46.69 -> 200
69.46.46.98 -> 200
69.46.46.110 -> 200

So the address your callback hostname happens to resolve to isn't special, and that gives you a test that separates "Railway is dropping this provider's requests" from "this one address is unreachable from their network". Ask Robokassa to retry against the same hostname pinned to a different address in 69.46.46.0/24, i.e. curl --resolve your-host:443:69.46.46.X https://your-host/payments/robokassa/result from the machine that actually sends the callbacks. If another address answers and yours doesn't, it's per-address reachability on the path between their network and that specific IP, and that is what Railway needs named in the ticket. There's an open thread right now about 69.46.46.41 being unreachable for one user, and another about misrouting for a specific origin AS, so it isn't a far-fetched failure mode.

While you're at it you can strike TLS off the list, because a .NET Framework/v4.0.30319 user agent is the first thing anyone will point at (that CLR leaves ServicePointManager.SecurityProtocol at an old default). Railway's edge floor is TLS 1.2, and it does present a fallback certificate when a client sends no SNI:

$ openssl s_client -connect docs.railway.com:443 -servername docs.railway.com \
    -tls1_1 -cipher 'ALL@SECLEVEL=0' -brief
... ssl3_read_bytes:tlsv1 alert protocol version ... SSL alert number 70

$ openssl s_client -connect docs.railway.com:443 -noservername -brief
Protocol version: TLSv1.3
Peer certificate: CN=*.up.railway.app

It doesn't fit your case: webhook.site enforces the same 1.2 floor (alert 70 there too) and their POST landed on it in 1ms, and a version mismatch returns an immediate alert rather than a 30-second hang. The edge also publishes no AAAA record, so there's no broken-IPv6 story to chase either. One caveat for later: with no SNI the edge answers as *.up.railway.app, which won't validate against a custom domain, so don't move the callback to a custom hostname on the assumption that clients send SNI.

To get payments flowing before any of that is resolved, put the callback endpoint on a custom domain proxied through Cloudflare. The provider then connects to Cloudflare's anycast instead of one address in Railway's /24, Cloudflare reaches your service from its own network, and you get a request log for the connections that do arrive, which is the evidence you're currently missing.

I can't tell you why that path is failing, and that part does need someone at Railway to pull flow logs for those four timestamps. But the retry-on-a-different-address test narrows it to a specific IP and gives them something concrete to look at.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...