Intermittent container DNS query loss, quantized to 5s/10.5s — one container 100% broken from birth passed healthcheck
thehimanshugarg
HOBBYOP

a month ago

Summary

A container deployed 2026-09-01 had broken outbound DNS from birth. Every outbound HTTPS call failed with Error: fetch failed after ~10.5s. It passed Railway's healthcheck (that path makes no outbound call), was promoted, and served 100% of traffic for 39 minutes, failing every authenticated request. A redeploy fixed it instantly. Independent corroboration: our APM agent went silent at 11:45:00 — container birth — and resumed the minute the replacement took over. Two unrelated outbound consumers failed and recovered together.

The failures are quantized, not scattered:

• Broken container: 10.45–10.50s then fetch failed (6 samples)

• Healthy container, intermittent 11:45–13:00 IST: 5.22–5.27s then success (7 of 881, ~0.8%)

• Healthy container, normal: 0.22–0.29s

The resolv.conf is search internal railway.internal + nameserver fd12::10, no options line — glibc defaults timeout:5, attempts:2. So 10.5s = both attempts lost; 5.25s = one lost. Integer multiples of the resolver timeout mean dropped queries, not a slow network. The upstream service answered our external probes in 0.42–0.85s during the same minutes, and /api/healthz (no outbound call) served 900+ consecutive 200s throughout.

The fault went quiet ~13:00 IST (780 clean samples in the 65 min after) — it is intermittent, so a spot check now will likely look clean. And it is per-query: at 13:00:57 our egress probe got 503 while another outbound call from the same container, in the same second, completed in 0.194s.

Environment: Trackr api / Trackr Core, Southeast Asia, Hobby, 1 replica, Dockerfile build, Node/Next.js, legacy environment (IPv6-only private DNS per your docs).

Questions:

(1) Is the resolver at fd12::10 known to drop queries? We measured ~0.8% for over an hour, and one container at 100% from birth.

(2) Can you inspect the host that ran deployment 38ff4f29-b603-4a0e-b866-162ed261f7b0 between 11:45 and 12:24 IST on 2026-09-01?

(3) Can we set resolv.conf options (shorter timeout) or a fallback nameserver? One dropped packet currently costs a 10-second user-facing stall.

(4) Could the legacy IPv6-only resolver path be the asymmetry? Is there an IPv4 fallback?

(5) Can the platform healthcheck detect dead egress? A container that cannot reach anything passes any probe that makes no outbound call — which is how this one was promoted.

Full timeline and ruled-out list in the first reply.

$10 Bounty

3 Replies

thehimanshugarg
HOBBYOP

a month ago

Full timeline (IST, UTC+5:30):

• 11:39:12 — deployment 38ff4f29-b603-4a0e-b866-162ed261f7b0 starts building

• 11:45:00 — its container goes live, broken from first breath

• 11:45:00 → 12:24 — 100% of outbound calls fail; Error: fetch failed dozens per minute in deploy logs

• 11:57:51 → 11:58:06 — first user-visible symptom (a login hang)

• 12:14:32 → 12:25:04 — independent 1Hz prober: 64 consecutive failures on the authenticated path

• 12:21:05 — manual redeploy (636c318a-9ff7-4ee1-ac94-577f07da2ad5)

• 12:25:07 — new container serving; recovery instant; 836 consecutive successes after

Also ruled out:

• Not the custom domain: identical 10.5s failures hitting the *.up.railway.app host directly.

• Not our code/config: same image, commit, and env before and after; only the container changed.

• Not resource pressure: node held 8 fds; /proc/net/sockstat showed TCP alloc 7909 / inuse 1, which reads as a consequence of thousands of 10s hanging attempts, not a cause.

• Not deploy churn: three older containers were removed mid-window; failures ran unchanged straight through their removal.

Post-recovery baseline: healthy egress to the same host is 35–92ms (median 51ms, 179 samples). The 5.25s stalls are ~100× that and land exactly on the resolver timeout.

Happy to provide raw logs and the full probe series.


Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


admilucky1-cyber
FREE

a month ago

Is fd12::10 known to drop queries?

Based on your measurements (~0.8% intermittent loss, one container at 100%), it looks like the resolver or its host path is indeed dropping packets. This isn’t normal behavior for glibc DNS resolution.

Inspecting host 38ff4f29… (11:45–12:24 IST)

Only Railway staff can check host logs. Your timeline and probe data are strong evidence that the host’s resolver was unhealthy during that window.

Can you set resolv.conf options or fallback?

By default, Railway doesn’t expose resolv.conf customization. You can try overriding DNS in your container (e.g., via docker run --dns=8.8.8.8) or by configuring Node’s resolver to use a custom nameserver. Shorter timeouts (e.g., options timeout:1 attempts:2) would reduce user‑visible stalls from 10s to ~2s.

IPv6‑only asymmetry

Yes, the lack of IPv4 fallback could be the culprit. If the IPv6 resolver path is flaky, you have no safety net. Forcing IPv4 DNS (Google 8.8.8.8, Cloudflare 1.1.1.1) may stabilize things.

Platform healthcheck detecting dead egress

Currently, Railway’s healthcheck only probes internal paths. A more robust check would include an outbound DNS+HTTPS probe. Without that, containers with broken egress can still be promoted.

✅ Practical Next Steps

Override DNS inside your container to a stable public resolver (Google or Cloudflare).

Reduce timeout in resolv.conf to avoid 10s stalls.

Raise with Railway support: Provide your probe logs and container IDs so they can inspect the resolver host.

Add your own egress healthcheck: e.g., a /api/egress endpoint that pings a known external host, so Railway’s probe can catch dead‑egress containers.


admilucky1-cyber
FREE

a month ago

Summary

On 2026‑09‑01, deployment 38ff4f29-b603-4a0e-b866-162ed261f7b0 in Southeast Asia (Trackr API / Trackr Core, Hobby plan, 1 replica, Node/Next.js, legacy environment) exhibited complete outbound DNS failure from container birth. Every HTTPS call failed with Error: fetch failed after ~10.5s. Redeploying (636c318a-9ff7-4ee1-ac94-577f07da2ad5) restored service instantly.

Evidence

Failure quantization:

10.5s stalls = both resolver attempts lost.

5.25s stalls = one attempt lost.

Normal baseline = 0.25s.

Broken container: 100% outbound failure from 11:45–12:24 IST.

Healthy container: ~0.8% intermittent stalls (5.25s) between 11:45–13:00 IST.

Independent probes: APM agent silent during failure window; resumed immediately after redeploy.

Healthcheck blind spot: /api/healthz (no outbound call) passed continuously, allowing promotion of a container with dead egress.

Upstream service: External probes answered in 0.42–0.85s during failures, ruling out remote slowness.

Timeline (IST, UTC+5:30)

11:39:12 — Build started.

11:45:00 — Container live, broken from first request.

11:45–12:24 — All outbound calls failed.

11:57:51 — First user‑visible login hang.

12:14–12:25 — Independent probe: 64 consecutive failures.

12:21:05 — Manual redeploy triggered.

12:25:07 — New container live; instant recovery; 836 consecutive successes.

Ruled Out

Custom domain (failures identical on *.up.railway.app).

Application code/config (same image/env before and after).

Resource pressure (low FD usage; sockstat consistent with hanging attempts).

Deploy churn (failures persisted across container removals).

Questions for Railway

Is the resolver at fd12::10 known to intermittently drop queries?

Can host logs for deployment 38ff4f29-b603-4a0e-b866-162ed261f7b0 (11:45–12:24 IST, 2026‑09‑01) be inspected?

Is there a supported way to set resolv.conf options (shorter timeout, fallback nameserver)?

Could the legacy IPv6‑only resolver path be the root cause? Is IPv4 fallback available?

Can platform healthchecks be extended to detect dead egress, preventing promotion of containers with broken outbound DNS?

Impact

39 minutes of complete outage for authenticated requests.

User‑visible login hangs.

Silent APM agent during failure window.

Redeploy resolved instantly, confirming host‑level fault.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...