Repeated Railway builder timeout reaching Docker Hub registry — 3/3 builds fail
celiomesan
FREEOP

a month ago

Hello Railway Support,

We are experiencing a reproducible build-infrastructure failure on our service.

Project: Centro de Consciência - Preview

Environment: preview

Service: centro-consciencia-platform

Region: EU West

Source: GitHub

Git commit: 6243be5794d43d2cd1936d8b98361a4bfad75928 (6243be5)

Commit message: chore(generative): add sanitized diagnostic observability

Dockerfile: apps/api/Dockerfile

We attempted the same deployment three times. All three attempts failed at the same build stage, before our application code was built.

The builder successfully loads apps/api/Dockerfile and then fails while resolving metadata for:

docker.io/library/node:22-alpine

The recurring error is:

failed to solve: DeadlineExceeded

node:22-alpine failed to resolve source metadata for docker.io/library/node:22-alpine

failed to do request: Head "https://registry-1.docker.io/v2/library/node/manifests/22-alpine"

dial tcp [...]:443: i/o timeout

Each failure occurs after approximately 30 seconds at the load metadata for docker.io/library/node:22-alpine stage.

Deployment attempts:

378a76c8 — FAILED with the same Docker Hub registry timeout.

c78c2c4e — FAILED. Deployment timestamp shown as 2026-08-28 00:43 GMT-3. Registry metadata request started around 00:43:33 and failed around 00:44:03.

815bba14 — FAILED. Deployment timestamp shown as 2026-08-28 00:54 GMT-3. Registry metadata request started around 00:55:01 and failed around 00:55:31.

Reproduction rate: 3/3 attempts.

We deliberately stopped after the third identical failure and have not initiated a fourth redeploy.

No code, Dockerfile, base-image reference, dependencies, environment variables, or deployment configuration were changed between these attempts.

Could you please investigate:

Was there a builder egress or registry connectivity problem affecting access from Railway's build infrastructure to registry-1.docker.io:443 during these deployments?

Were these three builds executed through the same builder pool, host, network path, NAT/egress route, or other shared infrastructure that could explain the repeated timeout?

Can you correlate the deployment IDs above with your internal builder/network telemetry and identify where the connection to Docker Hub timed out?

Can Railway requeue/rebuild this exact deployment on a healthy builder/pool/route, without requiring changes to our application, Dockerfile, node:22-alpine, dependencies, or environment configuration?

If this is a known or transient Railway-side issue, should we wait before attempting another build?

We can provide complete Build Logs and screenshots if required.

The application build itself never starts; all three attempts stop while resolving the node:22-alpine base-image metadata.

Railway's similar-discussion search surfaced a previous resolved case titled “Repeated Docker Hub metadata stall on Metal builder — 3 builds stuck on FROM node:22.12-slim”, where multiple builds also stalled while fetching Docker Hub base-image metadata and a Railway employee later noted that builds had resumed and that Railway would investigate why they had been stuck. Our current incident appears similar at the builder/registry boundary, although our failures terminate after approximately 30 seconds with DeadlineExceeded / i/o timeout.

Thank you.

$10 Bounty

1 Replies

Railway
BOT

a month ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 1 month ago


dolapobayo42-lgtm
FREE

a month ago

I can't confirm the root cause here since it requires visibility into Railway's builder egress path to registry-1.docker.io, which only Railway staff can check — but a couple of things worth doing while you wait on that:

This class of failure (builder → Docker Hub metadata resolution timeout) is a known recurring pattern — you referenced a similar prior thread yourself, which is a good sign it's a transient builder/egress issue rather than something in your Dockerfile.

As a workaround rather than a fix: try pinning to a digest instead of a tag (node:22-alpine@sha256:...) — sometimes tag resolution is what's timing out even when the underlying manifest is cached, and a digest pin can skip that lookup path. Also worth trying a rebuild with cache disabled, since a stuck/partial metadata fetch can sometimes get cached as a false failure.

The real fix here does need Railway to confirm whether your builds are landing on the same builder pool/route each time — that part isn't something outside telemetry can answer.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...