17 days ago
We’re troubleshooting an unusual WebRTC/RTP issue on Railway and have traced it as far upstream as Railway’s container permissions allow.
This is not a request to expose an inbound UDP port or run TURN/Coturn. Our Railway application initiates the WebRTC session, ICE and DTLS connect successfully, and both audio and video RTP reach the container.
The puzzle is that the simultaneous streams behave very differently:
Audio: 447/447 RTP sequence positions received — 0 apparent missing
H.264 video: 2,502/4,053 received — 1,551 apparent missing (38.3%)
Video had 1,031 sequence-gap events, with 0 duplicates and 0 late/out-of-order packets.
We initially suspected our Python/WebRTC stack, so we progressively instrumented the receive path upstream.
We now capture RTP immediately after Linux/Python socket.recvfrom(), before aioice/aiortc processing. The video sequence gaps are already present there.
Packets that do reach recvfrom() are preserved through aioice, ICE delivery, SRTP, RTP routing, and receiver/jitter-buffer input. Linux UDP receive-queue diagnostics also showed zero socket-queue drops.
We then tried to observe the traffic one step earlier using Linux AF_PACKET, but Railway correctly blocks this:
CAP_NET_RAW: false
CAP_NET_ADMIN: false
AF_PACKET: PermissionError(1, 'Operation not permitted')
Our question
Is there a Railway-supported way to determine whether these missing video UDP/RTP datagrams reached Railway before delivery to our container?
For example, could Railway provide or inspect host-side packet capture, eBPF/network telemetry, virtual-network drop counters, or another observation point upstream of the container?
We are not assuming Railway is dropping the packets. If they never reached Railway, that would be equally useful information.
We can run a fresh 30-second controlled test and provide the exact UTC window, Project/Service/Deployment IDs, RTP SSRC/payload type, and exact received/missing sequence numbers.
Essentially, we’ve reached recvfrom() from inside the container and need help moving the observation point one hop upstream.
Full diagnostic logs and instrumentation code are available if useful.
Pinned Solution
16 days ago
Thank you—this is very helpful. Network Flows would have given us exactly the additional Railway-side observation point we were looking for.
Since posting, we ran another controlled experiment that produced a pretty striking result. We moved the same application to AWS Lightsail, keeping the Ring camera, WHEP workflow, application code, aiortc 1.15.0 and diagnostic instrumentation essentially unchanged.
We have now run five independent AWS Live View sessions. All five showed 0% apparent missing H.264 sequence positions and all five successfully decoded the first 1280×720 frame. On Railway we had repeatedly seen roughly 40% apparent missing H.264 sequence positions and could not decode the first frame.
That doesn't prove precisely where the Railway-path loss occurred, but it changes our focus considerably. I'm going to pass your Network Flows suggestion to Railway support and ask whether they can correlate their platform-side observations with one of our affected UTC windows and missing-sequence lists.
Thanks again. Hopefully documenting this saves someone else from spending as much time as we did investigating H.264, NACK/RTX, keyframes and the decoder before testing the hosting/network path.
1 Replies
17 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 17 days ago
16 days ago
Thank you—this is very helpful. Network Flows would have given us exactly the additional Railway-side observation point we were looking for.
Since posting, we ran another controlled experiment that produced a pretty striking result. We moved the same application to AWS Lightsail, keeping the Ring camera, WHEP workflow, application code, aiortc 1.15.0 and diagnostic instrumentation essentially unchanged.
We have now run five independent AWS Live View sessions. All five showed 0% apparent missing H.264 sequence positions and all five successfully decoded the first 1280×720 frame. On Railway we had repeatedly seen roughly 40% apparent missing H.264 sequence positions and could not decode the first frame.
That doesn't prove precisely where the Railway-path loss occurred, but it changes our focus considerably. I'm going to pass your Network Flows suggestion to Railway support and ask whether they can correlate their platform-side observations with one of our affected UTC windows and missing-sequence lists.
Thanks again. Hopefully documenting this saves someone else from spending as much time as we did investigating H.264, NACK/RTX, keyframes and the decoder before testing the hosting/network path.
Status changed to Solved newcoderkc • 16 days ago
