2 months ago
Project: adaptable-luck (119455b3-fadc-470b-b192-c8e9ccc4fb9f)
Environment: production (5c37e9c5-52dc-4557-8177-2d6d585789da)
Service: airtag-tracker (80413133-4e85-4a39-8955-731ce77edf0c)
Deployment: fd54d81c-a938-46b4-847f-04ba18304b2f
Region: sfo
Timeline:
- Last successful location fetch: 2026-08-14 08:36:54 UTC
- First all-zero poll: 2026-08-14 08:40:21 UTC
- Failure remains active now.
The service remains healthy and runs its five-minute scheduler, but outbound
TCP traffic to gateway.icloud.com is dropped by Railway networking:
-
DNS queries for gateway.icloud.com succeed.
-
Recent network logs: 113 Apple-bound TCP/443 egress flows in 20 minutes;
all have dropCause=IP_RPFILTER.
-
The application then times out while fetching from Apple for every connected
account, so no locations or snapshots are saved.
Please investigate why this deployment’s egress traffic to the resolved
gateway.icloud.com IPs (17.248.193.x) is being dropped with IP_RPFILTER,
and advise whether the instance or its network path should be replaced.
I can provide exported network-log entries if useful.
3 Replies
2 months ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 2 months ago
2 months ago
Recovery update — I restarted airtag-tracker at 2026-08-15 13:47 UTC.
The service recovered immediately
The new instance’s Apple-bound traffic is no longer showing IP_RPFILTER. Before the restart, Railway network logs recorded 148 dropped TCP/443 flows to gateway.icloud.com; post-restart flows are completing without that drop cause.
This confirms the outage was tied to the previous instance/network path. Please investigate the IP_RPFILTER drops and confirm the durable platform-side resolution.
2 months ago
Additional symptom: Google OAuth was also affected.
Between 2026-08-15 13:30:13 and 13:32:15 UTC, five consecutive Google login attempts reached our callback but then returned users to the login page.
GET /auth/google/login completed normally (303 in 3–26 ms). After the user selected their Google account, /auth/google/callback took approximately 10 seconds (10,019–10,094 ms) before returning 303 to /login.
Application log for each attempt:
oauth google login failed: token exchange failed: [Errno 101] Network is unreachable
The failing operation is the server-side POST to:
https://oauth2.googleapis.com/token
which exchanges Google’s authorization code for tokens. No user session could be created, so the app redirected back to login.
This is not consistent with a bad OAuth redirect URI, client ID, or client secret: those failures should reach Google and return an HTTP error. Here the application failed locally while opening its outbound network connection.
After restarting airtag-tracker at 2026-08-15 13:47 UTC, a direct request from the new instance to the same Google token endpoint reached Google immediately (returning the expected HTTP 400 when sent without an authorization code). Google login connectivity therefore recovered with the restart.
Please correlate this Google OAuth egress failure with the prior instance’s Apple-bound TCP/443 flows dropped with IP_RPFILTER, and investigate the old instance/network path.
14 days ago
This really looks like an instance/network-path issue rather than anything specific to Apple.
IP_RPFILTER means the kernel dropped the packet because reverse-path validation failed. The part that makes this especially convincing is that Google OAuth egress was failing on the same instance too, then both Apple and Google started working immediately after the restart.
So I don't think I'd spend time changing DNS, Apple allowlists, OAuth config, etc. The common point is the old Railway instance/network path.
If it happens again, restarting/redeploying is probably the right immediate mitigation since that already moved you onto a working path.
For the actual root cause though, I think Railway would need to inspect the old instance/host. I'd give them the old deploymentInstanceId plus a few of the dropped flow IDs and timestamps.
You can also pull just the affected flows with something like:
railway logs --network \
--direction egress \
--drop-cause IP_RPFILTER \
--json \
--since 2026-08-14T08:30:00Z \
--until 2026-08-15T13:47:00ZThe useful fields there should be flowId, deploymentInstanceId, source/destination IPs and the capture timestamps.
Since the replacement instance immediately stopped producing IP_RPFILTER, I'd say the restart already confirmed the practical recovery path.
I don't think anyone outside Railway can really confirm the durable platform-side fix though, because that would require looking at the routing/rp_filter state on the old host.