18 days ago
Hello,
I have a staging-only service that repeatedly fails before the build process actually starts.
Project:
teraplace
Service:
teraplace-staging-migrator
Environment:
staging
Project ID:
9b398279-fd92-4a4c-85bb-0228b4c692af
Environment ID:
40b97d5f-e754-427e-a394-6963cbf16c0b
Service ID:
3c27fe46-ad3e-468d-88b1-6b3d7e781950
Failed deployments:
b8385977-d6af-4bd9-bf33-97290d32b422
c0f72b6d-6bc2-4737-81e1-7d8d1145c596
Both deployments fail immediately after:
scheduling build on Metal builder "builder-cwyigy"
There are no subsequent logs for:
- source checkout
- Node setup
- pnpm setup
- dependency installation
- build command
- container start
- migration command
The exact same failure was reproduced twice.
The repository and branch are valid and another service in the same
staging environment builds successfully from the same repository.
Working API service:
3c520952-fe91-4f0e-bba1-19373835bfcb
Repository:
iuri-maroto/teraplace
Branch:
main
The migration service is intentionally staging-only and has no public
domain.
No production resources are involved.
Could you please check whether these deployments are failing at the
Metal builder infrastructure/scheduler layer?
If this service is affected by a Metal builder issue, please advise
whether it can/should temporarily use the legacy/non-Metal builder,
or whether another remediation is recommended.
Please note that redeploy has already been attempted twice with the
same result, so I do not want to continue retrying without identifying
the cause.
Thank you.
5 Replies
18 days ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • 18 days ago
18 days ago
I found a recent Railway thread with a very similar visible symptom:
scheduling build on Metal builder ...
In that case, the public logs stopped at builder scheduling, but Railway
internally identified the actual failure at the BUILD_IMAGE step:
Root directory "backend" was not found in the deployed source.
Our case may have a different root cause, because this service uses
root / and another service in the same project successfully builds
from the same repository/branch.
Could you please inspect the internal BUILD_IMAGE failure reason for
these two deployments, rather than relying only on the public build logs?
b8385977-d6af-4bd9-bf33-97290d32b422
c0f72b6d-6bc2-4737-81e1-7d8d1145c596
The public logs for both stop immediately after:
scheduling build on Metal builder "builder-cwyigy"
We would specifically like to know whether Railway is rejecting the
deployment before checkout because of source/root/build configuration,
or whether this is actually a builder infrastructure issue.
We have stopped retrying until the underlying BUILD_IMAGE reason is known.
18 days ago
I found a few recent Railway incidents with the exact same visible symptom:
scheduling build on Metal builder ...
In those cases, the visible log was misleading because the actual failure
was an internal pre-build validation error, such as:
- Root Directory being resolved incorrectly;
- Root Directory being applied twice;
- an invalid Root Directory value;
- Railpack generating an empty secret ID from a malformed variable name.
For our service I have already verified:
Service:
teraplace-staging-migrator
Root Directory:
/
Source:
iuri-maroto/teraplace
Branch:
main
The working API service in the same staging environment also uses root /
and builds from the same repository/branch.
The migrator's custom variable name is:
MIGRATION_DATABASE_URL
and the exposed variable names do not show leading/trailing spaces or
an empty variable name.
Could you please inspect the internal BUILD_IMAGE failure for:
b8385977-d6af-4bd9-bf33-97290d32b422
c0f72b6d-6bc2-4737-81e1-7d8d1145c596
Specifically, could you check:
-
the resolved source/root directory used internally;
-
the generated Railpack build plan;
-
whether any build secret ID is empty or invalid;
-
the actual BUILD_IMAGE error hidden behind the public
scheduling build on Metal builderlog.
We have stopped retrying until the internal failure reason is identified.
18 days ago
Additional context after reviewing recent Railway incidents with the same
visible scheduling build on Metal builder symptom:
We have also ruled out the Root Directory issues seen in these cases:
-
The service is deployed directly from the GitHub integration, not via
railway upwith--path-as-root, so there should be no duplicatedroot-directory scoping.
-
The configured Root Directory is simply:
/
It is not a repository URL and does not point to a nested/nonexistent
folder.
-
The working API service uses the same repository, branch and root
/. -
The migrator's custom variable name exposed by Railway is exactly:
MIGRATION_DATABASE_URL
We do not see any empty variable name or leading/trailing whitespace.
This makes us suspect that another internal pre-build / BUILD_IMAGE validation
error is being hidden behind the public:
scheduling build on Metal builder "builder-cwyigy"
Could you please provide the actual internal failure reason for:
b8385977-d6af-4bd9-bf33-97290d32b422
c0f72b6d-6bc2-4737-81e1-7d8d1145c596
17 days ago
Great thread.
You ruled out essentially everything user-side, so there's little left but the internal BUILD_IMAGE reason. Two things while you wait on staff:
Workarounds worth trying: create a brand-new service from the same repo/branch (fresh service config), and/or toggle the builder setting off and on to force a new builder assignment — subtle service-config corruption won't show in your checks but a fresh service sidesteps it. Also worth a hard redeploy after clearing build cache.
Faster routing: for infra-side questions like this, the Railway Discord (discord.gg/railway) tends to get staff eyes much faster than the forum, especially with deployment IDs in hand like you have.