Repeated pre-Docker BUILD_IMAGE failure on the same Railway Metal builder
kandidcandid-spec
FREEOP

a month ago

Hello Railway team,

I am seeing repeated Worker image-build failures that occur before any Dockerfile instruction begins.

Workspace plan:

Hobby

Project:

khidyao-content-studio-staging

Environment:

staging

Service:

@khidyao/worker

Service ID:

1518ebdc-fb0e-426c-8a8b-8ee477104ce4

Region:

asia-southeast1-eqsg3a

Target commit:

6ec87a0e53810911ff0198800bdc5433e7d01cac

Failed deployment IDs and timestamps:

  1. eccce0dc-ab47-45a0-a9e4-aa34652612b3

    Created: 2026-08-25T23:58:27.260Z

  2. 2535a617-f805-4077-a45a-41483742f04c

    Created: 2026-08-26T04:12:21.930Z

    BUILD_IMAGE: 2026-08-26T04:12:23.691Z

             through 2026-08-26T04:12:29.288Z

    Terminal failure: 2026-08-26T04:12:29.305Z

Assigned Metal builder:

builder-nxrjjq

Observed behavior:

  • SNAPSHOT_CODE completes successfully.

  • BUILD_IMAGE begins and fails approximately five to six seconds later.

  • The complete public build log contains only:

    scheduling build on Metal builder "builder-nxrjjq"

  • The structured public error contains only:

    Failed to build an image. Please check the build logs for more details.

  • No evidence appears for:

    • Dockerfile parsing
    • FROM/base-image resolution
    • dependency installation
    • Worker compilation
    • pnpm deploy
    • final image assembly
    • image registry push
    • container creation
    • runtime startup
    • healthcheck execution
  • Multiple Worker deployments at different commits and times stopped after assignment to the same Metal builder.

  • A separate Migration service assigned to another Metal builder emitted normal BuildKit steps and completed successfully.

  • The resolved Worker deployment manifest correctly selects:

    • Dockerfile builder
    • Dockerfile.worker
    • repository-root build context
    • no custom build command
    • no custom start command
    • /health healthcheck
    • one Singapore replica
    • ON_FAILURE with three retries
  • Local Worker production build and deploy-tree validation pass, including:

    • dist/index.js production output
    • production dependencies
    • Node syntax and dependency resolution
    • non-root UID/GID 10001 runtime contract
    • /health and /ready route contracts

This does not appear to be an application-level build failure because no Dockerfile instruction is reached.

Could you please:

  1. Inspect the internal Metal builder/allocation failure for the deployment IDs and UTC timestamps above.
  2. Confirm whether builder-nxrjjq is unhealthy or has stale service-specific build state.
  3. Reassign a future build to another healthy builder, or advise how to clear the affected service/builder state without changing application source.
  4. Provide the hidden structured failure reason or internal error code, since the public logs contain no Dockerfile-level error.
  5. Confirm whether it is safe to create one new deployment after any platform-side correction.

We have stopped further retries and have not modified source or Worker configuration after reproducing the platform-side failure.

No credentials, secrets, connection URLs, customer content, or personal data are included in this report.

Thank you.

Solved

2 Replies

Railway
BOT

a month ago

Both deployments failed with the same Dockerfile validation error: dockerfile invalid: docker VOLUME at Line 25 is not supported, use Railway Volumes. This fires before the build reaches the Metal builder, which is why the only public log line visible after it is the builder scheduling message. Remove the VOLUME instruction from line 25 of your Dockerfile.worker and the build will proceed normally. If your service needs persistent storage, attach a Railway Volume to the service instead.


Status changed to Awaiting User Response Railway • about 1 month ago


kandidcandid-spec
FREEOP

a month ago

Thank you. Your hidden Dockerfile validation diagnosis was correct.

We confirmed and removed:

VOLUME ["/var/lib/khidyao/tmp"]

The path is ephemeral Worker scratch, while durable structured state remains in

PostgreSQL and durable artifacts remain in S3-compatible storage. No Railway

Volume was required.

We also added static regression coverage to prevent unsupported VOLUME

instructions from returning to Railway deployment Dockerfiles.

A bounded follow-up corrected the Worker production workspace-package closure.

The final image then built, pushed, and started successfully.

The remaining healthcheck mismatch was resolved by setting the Worker-specific:

PORT=3001

Final Worker deployment:

34990d3a-306e-42aa-b6b7-ccf31ba677db

Final result:

  • Deployment: SUCCESS
  • Instance: RUNNING
  • Region: Southeast Asia / Singapore
  • Runtime: node dist/index.js
  • Runtime UID/GID: 10001/10001
  • Listener: 0.0.0.0:3001
  • Railway healthcheck destination port: 3001
  • /health: HTTP 200
  • /ready: HTTP 200
  • Queue schema: ready, version 37
  • Storage: ready
  • Claims paused: true
  • Media capability: disabled
  • Providers: unavailable

We observed the idle Worker for more than 15 minutes. No application job,

claim, lease, retry, render, generation, publication, analytics run, provider

call, or Bucket mutation occurred.

This solved the deployment issue. Thank you for exposing the hidden validation

error.


Status changed to Awaiting Railway Response Railway • about 1 month ago


Status changed to Solved Railway • about 1 month ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...