something short like Deploys stuck at "scheduling build" on builder-gfasfb — 4 consecutive failures
malligaarjunr
FREEOP

3 months ago

Subject: Deploys stuck at "scheduling build" on builder-gfasfb — 4 consecutive failures, no build/deploy logs

Project: ctrak (project ID: 55233d5a-eba1-483e-90a8-79ef9cfb89b9)

Service: backend (service ID: bb3fd1ac-e73c-44c7-81f5-2fae1058dcb2)

Summary:

Four consecutive railway up deployments to our "backend" service have

failed, all with an identical symptom: the build log contains only a

single line — "scheduling build on Metal builder 'builder-gfasfb'" —

and never progresses to an actual build step, let alone a deploy step.

No error message, no stack trace, just silence after that line until

the deployment is marked FAILED.

Failed deployments (all same symptom, same builder ID):

  1. 3059a574-1c02-4ca5-a9be-072400fc6b29 — FAILED — 2026-07-13 13:14:49 -04:00

  2. 4c7c7dbd-c234-4422-a77f-2683ff78a4b0 — FAILED — 2026-07-13 13:15:34 -04:00

    (retried immediately after #1, same result)

  3. 31d7965b-c960-4450-a318-a851373dff6c — FAILED — 2026-07-13 13:21:16 -04:00

    (retried after a 2.5-minute wait, same result)

  4. 234f47f0-053b-4a09-8020-46cd9d1ca742 — FAILED — 2026-07-13 13:37:44 -04:00

    (retried from the repo root instead of the backend/ subdirectory,

    specifically to rule out a working-directory/Root Directory

    mismatch — same result)

Notable correlation:

The same builder ID ("builder-gfasfb") also appears on an older failed

deployment from several days earlier:

3018e9d7-f7ae-4a94-a759-bdd11f2c60ee — FAILED — 2026-07-09 00:06:25 -04:00

That earlier failure was tied to a separate, since-resolved issue

(an unintended GitHub auto-deploy connection on this service, which we

have since disconnected via railway service source disconnect). The

recurrence of the exact same builder ID across unrelated deploy

attempts, days apart, suggests this specific builder instance may be

unhealthy rather than an issue on our end — our code has not changed

between the deploy that succeeded and the four that failed.

What we've already ruled out:

  • Not a code/build config issue: the same source successfully deployed

    earlier today (deployment 92645a79-f9c8-4df3-954f-60763c6936f4,

    SUCCESS, 2026-07-13 12:00:39 -04:00) and is still running now.

  • Not our GitHub integration: confirmed via railway service list --json

    that the service's source is null (cleanly disconnected).

  • Not Root Directory / working directory: confirmed in the dashboard

    that Root Directory is correctly set to "backend". As an additional

    test, we ran railway up --service backend from the repo root

    instead of the backend/ subdirectory (deployment 234f47f0 above) to

    rule out any working-directory mismatch entirely — it failed with

    the identical symptom, same builder ID. The common factor across all

    four failures is specifically builder-gfasfb, regardless of source

    path, timing, or retry gap.

  • Not a reported platform-wide incident: checked status.railway.com,

    which shows "Fully Operational," no current incidents.

  • Not a transient blip: retried four times total, including once after

    a 2.5-minute wait and once from a different working directory, with

    identical results each time.

Solved

14 Replies

Status changed to Awaiting Railway Response Railway • 3 months ago


sam-a
EMPLOYEE

3 months ago

Builder-gfasfb itself is healthy, it's building other services normally. What's happening is that builds for your backend service fail a couple of seconds after scheduling, before any output is produced, which is why the log never gets past that first line. Builder assignment is sticky per service, so every retry lands on the same builder no matter where you run railway up from. We're tracking this pattern with other reports.

The fix you can do yourself: delete and recreate the backend service with the same repo and settings. A fresh service gets a fresh builder assignment. There's no volume on this service so no data is at risk, just copy your variables over first.

Railway Team


Status changed to Awaiting User Response Railway • 3 months ago


malligaarjunr
FREEOP

3 months ago

Hi,

Following up on your suggested fix (delete and recreate the service so it gets a fresh builder assignment). We tried that, plus an additional region variation, and the exact same failure reproduced both times — this looks broader than a single stuck builder assignment on our original service.

Project: ctrak (55233d5a-eba1-483e-90a8-79ef9cfb89b9), environment production (adce65f7-3742-42b6-8749-4a45f3726013).

For context, here's the original issue: service "backend", region sfo, builder "builder-gfasfb" — stuck at "scheduling build," never progressed, across 4 attempts last night.

Tonight we applied your suggested fix: deleted the original service entirely and created a brand-new one, "backend" (bb692f59-630c-40ca-b9e6-15a81dbf7281), no GitHub source connected, default region sfo, all env vars restored, single railway up attempt. Same symptom, just a different builder pod this time. Deployment 985afac2-c0a8-4c79-9423-22c1ea83b838, and the build log shows nothing beyond two identical lines — "scheduling build on Metal builder builder-pipjvr" at 17:41:48 and again at 17:41:51 — before the deployment flipped straight to FAILED with no build steps ever starting.

To rule out a region-specific builder pool issue, we repeated the same process once more with the region explicitly set to us-west2 instead of sfo: another brand-new service, "backend" (d907bb0d-19dd-4974-b048-004625c2d552), same env vars restored, single deploy attempt. Same result again, on yet another builder pod. Deployment e18e7371-0e6e-4b07-9508-b61318ce5a98, build log again just two lines — "scheduling build on Metal builder builder-qniiyc" at 18:03:38 and 18:03:43 — then FAILED with nothing further.

So across three separate services, three separate builder pod assignments, and two separate regions, we're seeing the identical failure. That rules out both "this one service had a bad builder assignment" and "this is a region-specific builder pool problem" — it looks like something's wrong at a broader level on your side, account/project/builder-pool-wide, not something we can route around by recreating services ourselves.

We haven't attempted any further retries beyond these three data points and won't until we hear back. Happy to provide any additional logs or IDs you need. This has now blocked our production auth cutover for over 24 hours, so a fast turnaround would be really appreciated.

Thanks.


Status changed to Awaiting Railway Response Railway • 3 months ago


3 months ago

Your testing was thorough, and the recreate-service suggestion we gave earlier was not the right fix for what you're experiencing. Cancel any stuck or queued builds on the current service, wait about 15 minutes, then try a fresh railway up. If the deploy still gets stuck at the scheduling step, reply here with the new deployment ID.


Status changed to Awaiting User Response Railway • 3 months ago


malligaarjunr
FREEOP

3 months ago

Hi noahd,

Followed your steps exactly: checked for stuck or queued builds first (none found — both prior deployments were already terminal, nothing to cancel), waited the full 15 minutes, then ran a fresh railway up.

Result: same symptom. New deployment ID: 1f66c5f6-3ae5-4a8c-ad2b-4075c431d344

One more thing worth noting: this new attempt landed on builder-qniiyc — the same builder pod as one of our three data points from yesterday. That's now 2 of 4 total attempts landing on that specific builder. Might be worth checking that pod specifically rather than the account/project level.

Thanks.


Status changed to Awaiting Railway Response Railway • 3 months ago


sam-a
EMPLOYEE

3 months ago

Thanks, that's helpful. The scheduling step is actually succeeding here, a builder gets picked fine. The failure is happening a few seconds later, before the build produces any output, and it's repeating on every builder, so we want to catch a fresh attempt while we can still see the details on our side.

Could you run one more deploy straight from your terminal with railway up --verbose (not through any editor or agent integration), and send us:

  • the new deployment ID
  • the time you ran it, with timezone
  • the full CLI output

The verbose flag may surface more about what's failing, and running it directly from the terminal rules out anything on the upload side. The deployment ID and timestamp are the main things, they let us pull the internal build record before it ages out. Once we have that we'll know whether this needs to go to our builds team.

Railway Team


Status changed to Awaiting User Response Railway • 3 months ago


sam-a

Thanks, that's helpful. The scheduling step is actually succeeding here, a builder gets picked fine. The failure is happening a few seconds later, before the build produces any output, and it's repeating on every builder, so we want to catch a fresh attempt while we can still see the details on our side. Could you run one more deploy straight from your terminal with `railway up --verbose` (not through any editor or agent integration), and send us: - the new deployment ID - the time you ran it, with timezone - the full CLI output The verbose flag may surface more about what's failing, and running it directly from the terminal rules out anything on the upload side. The deployment ID and timestamp are the main things, they let us pull the internal build record before it ages out. Once we have that we'll know whether this needs to go to our builds team. Railway Team

malligaarjunr
FREEOP

3 months ago

Ran it directly from a plain terminal, not through any editor/agent integration, per your request.

New deployment ID: 1b87ca55-1206-4992-81e0-85dc3cece573

Time run: 2026-07-15 19hr 01min

(UTC-05:00) Eastern Time (US & Canada)

Full CLI output:

PS C:\Users\malli\Projects\construction-platform\backend> railway up --verbose

New version available: v5.26.1 visit https://docs.railway.com/guides/cli for more info

Indexed

Compressed [====================] 100%

railway up

service: d907bb0d-19dd-4974-b048-004625c2d552

environment: adce65f7-3742-42b6-8749-4a45f3726013

bytes: 226225

Uploaded

Build Logs: https://railway.com/project/55233d5a-eba1-483e-90a8-79ef9cfb89b9/service/d907bb0d-19dd-4974-b048-004625c2d552?id=1b87ca55-1206-4992-81e0-85dc3cece573&

Deploy failed


Status changed to Awaiting Railway Response Railway • 3 months ago


sam-a
EMPLOYEE

3 months ago

We found where it's failing. The scheduling step is fine, the failure happens in the build planner, which analyzes your code and exits with an error about two seconds in, on every attempt. That's why it reproduced across services, regions, and builders. Separately, the planner's error message isn't making it into your build log, which is a bug on our side and why this looked so opaque.

The most likely trigger: your recreated backend service has no Start Command set, and the planner errors out when it can't determine one on its own. Your original service had more config on it, which would explain why the same code built fine there before.

Could you set an explicit Start Command on the backend service (Settings, Deploy section), for example your java -jar command, plus a Build Command if you use one, then run railway up again? If it still fails with that set, send us the new deployment ID and we'll take it to our builds team.

Railway Team


Status changed to Awaiting User Response Railway • 3 months ago


sam-a

We found where it's failing. The scheduling step is fine, the failure happens in the build planner, which analyzes your code and exits with an error about two seconds in, on every attempt. That's why it reproduced across services, regions, and builders. Separately, the planner's error message isn't making it into your build log, which is a bug on our side and why this looked so opaque. The most likely trigger: your recreated backend service has no Start Command set, and the planner errors out when it can't determine one on its own. Your original service had more config on it, which would explain why the same code built fine there before. Could you set an explicit Start Command on the backend service (Settings, Deploy section), for example your java -jar command, plus a Build Command if you use one, then run railway up again? If it still fails with that set, send us the new deployment ID and we'll take it to our builds team. Railway Team

malligaarjunr
FREEOP

3 months ago

Hi,

Set the Custom Start Command as you suggested (java -jar $(ls build/libs/*.jar | grep -v plain), matching our original service's config) and ran a fresh railway up.

Result: same symptom. New deployment ID: 8d694b29-2333-49c9-9a3f-c8da8763d38b

Build log again just two lines, same shape as every prior attempt:

2026-07-16T03:49:18.976Z scheduling build on Metal builder "builder-qniiyc"

2026-07-16T03:49:21.893Z scheduling build on Metal builder "builder-qniiyc"

This rules out the Custom Start Command theory — the failure happens before the build planner would ever reach that config, so an empty Start Command wasn't the trigger.

One more data point that may help narrow this down: this is now the THIRD attempt (out of 5 total, across two regions) landing specifically on builder-qniiyc. Given that pod-level recurrence, would it be worth checking that specific builder instance directly, or is it time to take this to your builds team as you mentioned earlier?

Thanks for staying on this with us.


Status changed to Awaiting Railway Response Railway • 3 months ago


malligaarjunr
FREEOP

3 months ago

Hi,

Just checking in — any update, or has this been passed to your builds team yet? We're now several days into this blocking our production deployment, so any status update would help us plan around it.

Thanks again for digging into this with us.


sam-a
EMPLOYEE

3 months ago

We dug into the internal records for every attempt and found the concrete difference. Your one successful deploy included a railway.json file in the upload, which set the start command. None of the failed uploads had that file, and the start command on the current backend service is still empty on our side, so the value you set in the dashboard never saved to this service. That's easy to do here since you've had more than one service named backend during the recreates.

Two ways to fix it, either works:

  1. Add a railway.json next to your code with {"deploy": {"startCommand": "java -jar $(ls build/libs/*.jar | grep -v plain)"}} and run railway up from that directory, or
  2. Open the backend service in the us-west2 project, set the Custom Start Command in Settings, and confirm it shows as saved when you reload the page, then deploy.

The missing error message in your build log is a real bug on our side. Sorry it made this so hard to see.

Railway Team


Status changed to Awaiting User Response Railway • 3 months ago


malligaarjunr
FREEOP

3 months ago

Hi sam-a,

I checked backend/railway.json before making any change, and I want to flag something before we go further: the file was already there, with the exact content you suggested, the entire time. Nothing needed fixing.

git log --follow -- backend/railway.json:

9030de9 2026-07-08 23:48:27 -0400 Fix railway.json start command: exclude plain jar from glob match

fa088f5 2026-07-08 23:44:18 -0400 Add railway.json: explicit start command for Railway deploy

Both commits are from July 8th — days before last night's first deploy attempt and every attempt since. git diff and git status on the file show zero changes, ever. Its content has been:

{

"$schema": "https://railway.com/railway.schema.json",

"deploy": {

  "startCommand": "java -jar $(ls build/libs/*.jar | grep -v plain)"

}

}

That's functionally identical to what you asked us to set — the only difference is the optional $schema key, which is editor validation only and has no effect on deploy behavior. Every one of our 6 deploy attempts ran from the backend/ directory, so this file was included in the uploaded source every single time.

I want to be straightforward: this is the second suggested fix in a row that hasn't matched what we're actually seeing. The first was recreating the service (assuming a bad builder assignment) — we did that, twice, in two different regions, and got the identical failure both times. Now this one — a missing start command — turns out to have never been true either.

I'd also gently push back on whether a missing start command could cause this symptom at all, structurally. Every single failure across all 6 attempts happens at "scheduling build on Metal builder ...", with zero further log output, before Railpack ever starts analyzing our code. A start command is read later in the build/deploy pipeline, well after that point — it can't be the cause of a hang that happens before the build even begins.

One thing that has held up, and gotten stronger rather than weaker as we've gathered more data: 4 of our 6 total deploy attempts have landed on the same specific builder pod, builder-qniiyc. That's the one lead in this whole investigation that evidence actually supports.

Given that two proposed causes have now been ruled out directly by evidence, could this get escalated to the builds team now? I believe that was mentioned earlier as the next step if the start-command theory didn't resolve it. We really appreciate the time you've put into this — genuinely, thank you — we just want to make sure this doesn't keep cycling through more surface-level fixes when the data points pretty clearly at the builder infrastructure itself.

Thanks.


Status changed to Awaiting Railway Response Railway • 3 months ago


sam-a
EMPLOYEE

2 months ago

Sorry for the back and forth on this one. It turns out there is a bug on our side: the build planner's error never makes it into your build log, so every attempt looked like it died silently at scheduling.

The actual cause: your uploads contain the app under a backend/ folder, but the recreated service has no Root Directory set. Our build system only reads railway.json from the configured root, so it skipped backend/railway.json on every attempt, and with no start command reaching the planner it exited a couple of seconds in, on any builder, in any region. Your earlier point that a start command couldn't matter before the build begins was a fair read of the logs, but the planner does actually run about two seconds after the scheduling line, its output just never reached your log because of the bug above. The repeated builder-qniiyc assignment is expected behavior, builds for one service intentionally stick to the same builder.

To fix it: on the backend service in the us-west2 project, set Root Directory to backend in Settings, then deploy from the repo root. One more thing we spotted, the Custom Start Command you saved earlier never persisted on this service, it still shows empty on our side, so please set that again as well.

If the next deploy doesn't go through, send us the deployment ID within a day or two and we'll have full internal detail on it.


Status changed to Awaiting User Response Railway • 3 months ago


malligaarjunr
FREEOP

2 months ago

Hi,

Thanks for tracking this down — genuinely appreciate the persistence, and glad it turned out to be something concrete rather than an unexplained infra mystery. Good to know the builder-qniiyc recurrence was expected/benign too.

We ended up migrating the backend off Railway to another host a few days ago while this was still open, so we won't be testing the Root Directory / Custom Start Command fix on our end. Wanted to close the loop rather than leave you waiting on a deployment ID that's not coming.

Really appreciate the help throughout — this'll likely save someone else a few days if they hit the same "recreated service, missing Root Directory, silent planner error" combination.

Thanks again.


Status changed to Awaiting Railway Response Railway • 3 months ago


sam-a
EMPLOYEE

2 months ago

Thanks for closing the loop, and sorry this took as long as it did. We hope the migration went smoothly, and you're welcome back any time.


Status changed to Awaiting User Response Railway • 3 months ago


Status changed to Solved sam-a • 3 months ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...