Error
pozhovanii
FREEOP

16 days ago

Hello. I’m having a deployment issue with one Railway service.

Project: eth-agent-v1

Service: eth-agent-v1

Service ID: f59daba4-930b-47cb-a91d-219d3ab80a06

Environment ID: 7389b652-ad32-4eab-b174-bad9d059c9a8

The last successful deployment was: b6251d97-93ba-40f9-9e0f-4c60dcf3493f commit: f0ce81d3f0d7dca74ac4636da25a2b1a184b67f7

Every deployment after that fails before the actual Docker build starts.

The build logs contain only:

scheduling build on Metal builder "builder-geiqef"

and then the deployment changes to FAILED. There is no Dockerfile output, no pip output, no pytest output, and no application error.

The latest failed deployment is: 3c21340e-5091-45ff-9a9e-cfdf6c5818eb

I already tried redeploying and removing the custom build command, but the result is the same.

The same builder-geiqef successfully built this service immediately before the problem started. Also, another project in the same Railway workspace is currently able to build successfully on a different Metal builder, so this does not appear to be a workspace-wide build outage.

Could you please check whether this service is stuck on a problematic Metal builder or whether there is some internal builder/service state that needs to be reset or reassigned?

Please do not delete or recreate the service, database, environment variables, or the currently running last-known-good deployment.

Thank you.

$10 Bounty

1 Replies

Railway
BOT

16 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 16 days ago


linkindadaba
HOBBYTop 10% Contributor

15 days ago

Yes, your service is stuck on the Metal builder builder-geiqef, which has entered a degraded/wedged state (it accepts scheduling from the orchestrator but crashes before executing the source checkout or Docker build).

Why this is happening:

  1. Corrupted Builder Node: As you noticed, other projects work because they were assigned to different Metal builders. builder-geiqef itself is wedged.
  2. Builder Affinity Lock: Railway's build scheduler hashes service IDs to specific builder nodes to maximize Docker layer cache hits. Because of this affinity, every standard "Redeploy" or new commit on eth-agent-v1 gets routed straight back to the same dead node (builder-geiqef).

Non-Destructive Solutions (Keeps your service, database, and running deployment completely intact):

Option 1: Toggle off "Use Metal Build Environment" (Immediate Fix)

This safely reroutes your build away from the broken Metal builder pool to Railway's legacy build infrastructure:

  1. Open your eth-agent-v1 service.
  2. Go to Settings > scroll down to the Build section.
  3. Toggle "Use Metal Build Environment" to OFF (disabled).
  4. Trigger a new deployment.

Your currently running deployment (b6251d97...) will continue serving traffic uninterrupted until the new build succeeds.


Option 2: Force the Scheduler to Reassign a New Metal Builder

If you want to stay on Metal builders, you must break the scheduler's builder affinity:

  1. In your service's Deployments tab, click the three dots (...) on your latest commit and select "Deploy with No Cache".
  2. If that still hits builder-geiqef, go to Variables and add a temporary build variable (e.g. FORCE_REBUILD=1). Adding or changing an environment variable invalidates the builder's local cache key and often forces the scheduler to allocate an alternate, healthy builder node in the cluster.

Option 3: Check Git Diff against f0ce81d

Since commit f0ce81d3 built successfully on builder-geiqef right before the failure:

  • Check your git diff f0ce81d..HEAD to see if a file like railway.json, railway.toml, .dockerignore, or a root configuration was added or modified. A malformed config file or a typo in a newly added configuration can cause the builder's pre-flight parser to fail before emitting logs.

Note for Railway Staff:

The builder node builder-geiqef is currently wedged/dropping builds immediately after scheduling and needs to be recycled or pulled from the rotation pool. (Service ID: f59daba4-930b-47cb-a91d-219d3ab80a06, Deployment ID: 3c21340e-5091-45ff-9a9e-cfdf6c5818eb).


Welcome!

Sign in to your Railway account to join the conversation.

Loading...