2 months ago
Summary: Our service workout-tracker (project devoted-rejoicing / production) cannot connect to our MongoDB Atlas M0 cluster. Connections silently hang for the full timeout (no SYN-ACK, no RST) on port 27017, while all other outbound traffic from the same container works normally. Failing continuously since 2026-08-01 ~21:19 UTC.
What we already ruled out:
- Atlas IP Access List: 0.0.0.0/0 is active.
-
- Cluster health: Atlas shows "ready to use" (green), replica set healthy (1 PRIMARY + 2 SECONDARY).
-
- Connection string unchanged, standard mongodb+srv:// scheme.
-
- Both Railway and MongoDB Atlas status pages show fully operational.
-
- Restarted the deployment: no change.
-
- Moved the service region from US West to US East to match the Atlas cluster's stated region and redeployed: no change, identical NetworkTimeout errors immediately after the new deployment went active.
Diagnostic run directly inside the container via the Railway web Console:
Container egress IP: 162.220.234.6
DNS resolves fine to 3 distinct hosts:
customer-apps-shard-00-00.f03lwe.mongodb.net -> 35.238.151.175
customer-apps-shard-00-01.f03lwe.mongodb.net -> 35.239.145.160
customer-apps-shard-00-02.f03lwe.mongodb.net -> 34.41.217.74
General outbound works fine:
curl -s -o /dev/null -w "%{http_code}\n" --max-time 5 https://www.google.com
-> 200
Connecting to the Atlas cluster IP on port 27017 hangs for the entire timeout, no RST, no response at all (silent drop, not "connection refused"):
time (timeout 6 bash -c 'cat < /dev/null > /dev/tcp/35.238.151.175/27017')
-> real 0m6.002s, EXIT:124 (killed by timeout, never connected)
Control test, same port 27017, different destination (portquiz.net, built to accept connections on every port) connects instantly:
time (timeout 6 bash -c 'cat < /dev/null > /dev/tcp/portquiz.net/27017')
-> real 0m0.090s, EXIT:0
This rules out a blanket port-27017 block on Railway's side in general — arbitrary destinations on port 27017 work fine. The block is specific to the 3 Atlas cluster IPs (all Google-Cloud-owned addresses, consistent with the cluster's stated US East / Virginia region).
App-level symptom for context: our Python/PyMongo backend retries 3x with socketTimeoutMS=8000 and connectTimeoutMS=8000, every attempt fails with NetworkTimeout to all 3 shard hosts, resulting in a 503 on /api/auth/register and /api/auth/login. This is currently blocking all new user registrations and logins.
Ask: could you check whether egress IP 162.220.234.6 (or the underlying NAT range for this service) is being filtered/rate-limited on the way to Google-Cloud-hosted destinations, or whether reassigning/rotating the service's egress IP would help? Happy to provide any additional deployment/project details needed. Thanks for any help.
1 Replies
2 months ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 2 months ago
2 months ago
The portquiz.net:27017 control is useful: it rules out a blanket outbound-port block. A silent SYN timeout to all three Atlas hosts still has two plausible causes, though: Atlas is dropping this source address at its network boundary, or the path between this Railway egress range and those Google Cloud destinations is broken.
PyMongo credentials, retry counts, and TLS settings are not involved until the TCP handshake succeeds. Use the real Atlas hostnames and test the TLS boundary without putting the MongoDB URI in a command:
for host in \
'<atlas-host-00>' \
'<atlas-host-01>' \
'<atlas-host-02>'
do
printf '\n== %s ==\n' "$host"
getent ahostsv4 "$host"
timeout 10 openssl s_client \
-connect "$host:27017" \
-servername "$host" \
-brief </dev/null
doneRun the same openssl test from a laptop or another provider. A visible Atlas certificate or completed TLS handshake is enough for this test; no database password is needed.
- If the laptop also times out, the problem is on the Atlas cluster/network-access side. Confirm that
0.0.0.0/0is Active in the same Atlas project that owns this exact SRV hostname, not pending, temporary/expired, or configured in another project. - If the laptop succeeds but Railway times out, the failure is source/path-specific. The cleanest test on a Pro workspace is to enable Railway Static Outbound IPs, redeploy, add that exact IPv4 as a
/32Atlas IP access-list entry, wait until Atlas marks it Active, and repeat the hostname/TLS test.
Using a specific /32 is important here. It both changes the Railway source address and proves that the rule is attached to the correct Atlas project. Keep the existing rule until the new /32 is Active so you do not accidentally lock out the application. After the issue is resolved, remove the broad 0.0.0.0/0 entry rather than leaving the database open to every source address.
Interpret the static-IP test as follows:
- New static IP plus an Active
/32works: keep that pairing. The previous shared egress address or route was the problem. - The new
/32still times out but the laptop works: provide Railway and Atlas with the old/new source IPs, destination hostname/IP, deployment region, and one exact UTC test window. That is a platform routing/filtering incident rather than a PyMongo bug. - Both laptop and Railway fail: use Atlas's connection dialog/activity information to correct the cluster or project network-access configuration before changing Railway again.
MongoDB documents that Atlas accepts public connections only from addresses in the project IP access list and that clients must reach TCP ports 27015-27017. Railway documents static outbound IPv4 specifically for integrations such as MongoDB Atlas:
- Atlas IP access lists: https://www.mongodb.com/docs/atlas/security/add-ip-address-to-list/
- Atlas connection and port requirements: https://www.mongodb.com/docs/atlas/connect-to-database-deployment/
- Railway static outbound IPs: https://docs.railway.com/networking/static-outbound-ips
Do not post the full SRV URI, username, password, or TLS transcript if it contains account-specific names. Hostnames, resolved IPs, source IP, region, and timestamps are sufficient for network correlation.