16 days ago
Hello. I’m having a deployment issue with one Railway service.
Project: eth-agent-v1
Service: eth-agent-v1
Service ID: f59daba4-930b-47cb-a91d-219d3ab80a06
Environment ID: 7389b652-ad32-4eab-b174-bad9d059c9a8
The last successful deployment was: b6251d97-93ba-40f9-9e0f-4c60dcf3493f commit: f0ce81d3f0d7dca74ac4636da25a2b1a184b67f7
Every deployment after that fails before the actual Docker build starts.
The build logs contain only:
scheduling build on Metal builder "builder-geiqef"
and then the deployment changes to FAILED. There is no Dockerfile output, no pip output, no pytest output, and no application error.
The latest failed deployment is: 3c21340e-5091-45ff-9a9e-cfdf6c5818eb
I already tried redeploying and removing the custom build command, but the result is the same.
The same builder-geiqef successfully built this service immediately before the problem started. Also, another project in the same Railway workspace is currently able to build successfully on a different Metal builder, so this does not appear to be a workspace-wide build outage.
Could you please check whether this service is stuck on a problematic Metal builder or whether there is some internal builder/service state that needs to be reset or reassigned?
Please do not delete or recreate the service, database, environment variables, or the currently running last-known-good deployment.
Thank you.
1 Replies
16 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 16 days ago
15 days ago
Yes, your service is stuck on the Metal builder builder-geiqef, which has entered a degraded/wedged state (it accepts scheduling from the orchestrator but crashes before executing the source checkout or Docker build).
Why this is happening:
- Corrupted Builder Node: As you noticed, other projects work because they were assigned to different Metal builders.
builder-geiqefitself is wedged. - Builder Affinity Lock: Railway's build scheduler hashes service IDs to specific builder nodes to maximize Docker layer cache hits. Because of this affinity, every standard "Redeploy" or new commit on
eth-agent-v1gets routed straight back to the same dead node (builder-geiqef).
Non-Destructive Solutions (Keeps your service, database, and running deployment completely intact):
Option 1: Toggle off "Use Metal Build Environment" (Immediate Fix)
This safely reroutes your build away from the broken Metal builder pool to Railway's legacy build infrastructure:
- Open your
eth-agent-v1service. - Go to Settings > scroll down to the Build section.
- Toggle "Use Metal Build Environment" to OFF (disabled).
- Trigger a new deployment.
Your currently running deployment (b6251d97...) will continue serving traffic uninterrupted until the new build succeeds.
Option 2: Force the Scheduler to Reassign a New Metal Builder
If you want to stay on Metal builders, you must break the scheduler's builder affinity:
- In your service's Deployments tab, click the three dots (
...) on your latest commit and select "Deploy with No Cache". - If that still hits
builder-geiqef, go to Variables and add a temporary build variable (e.g.FORCE_REBUILD=1). Adding or changing an environment variable invalidates the builder's local cache key and often forces the scheduler to allocate an alternate, healthy builder node in the cluster.
Option 3: Check Git Diff against f0ce81d
Since commit f0ce81d3 built successfully on builder-geiqef right before the failure:
- Check your
git diff f0ce81d..HEADto see if a file likerailway.json,railway.toml,.dockerignore, or a root configuration was added or modified. A malformed config file or a typo in a newly added configuration can cause the builder's pre-flight parser to fail before emitting logs.
Note for Railway Staff:
The builder node builder-geiqef is currently wedged/dropping builds immediately after scheduling and needs to be recycled or pulled from the rotation pool. (Service ID: f59daba4-930b-47cb-a91d-219d3ab80a06, Deployment ID: 3c21340e-5091-45ff-9a9e-cfdf6c5818eb).