2 months ago
Since 2026-08-17 06:14 UTC, every deployment on both services in project ingenious-hope
(services titan-workers and titan-web-design, environment 08c85767...) fails to build.
The symptom is identical every time. Each deployment produces exactly ONE build log line and then nothing further:
2026-08-17T17:12:08.197060314Z scheduling build on Metal builder "builder-fstdwd"There is no subsequent build output, no image, no error. Deployment logs are empty (0 lines). Earlier attempts today eventually flipped to FAILED with that same single line as their entire log; the two most recent have now sat in BUILDING for over 20 minutes with the same single line.
Recent deployments, both services (identical pattern on each):
| Created (UTC) | Status | Commit |
|---|---|---|
| 17:12:04 | BUILDING (20+ min, one log line) | 1d85f6c0 |
| 17:05:40 | BUILDING (20+ min, one log line) | e89cdee5 |
| 13:55:37 | FAILED | f8ded24b |
| 13:17:00 | FAILED | 2f995124 |
| 12:52:49 | FAILED | e7c5739b |
| 11:17:27 | FAILED | 1ac0e3fe |
| 10:54:45 | FAILED | b455bbf3 |
Two example wedged deployment IDs: c3e6e779... (titan-workers) and a5a8e009...
(titan-web-design).
What we have already tried:
- Reinstalled the Railway GitHub App and re-bound both services to the repository. This DID fix a separate earlier problem ("repository not authorized") - deployments are now created correctly from the right commits, so the GitHub trigger side is working. The builds themselves still never leave the scheduling step.
- Cancel + redeploy via the public API, three times. Each attempt was refused with a message to wait for the original build to finish - but the original never finishes.
- Pushed a fresh empty commit as a clean retrigger. Same single log line.
Your status page reported operational throughout.
What we need: the wedged builds aborted on your side, and the Metal builder assignment for these two services unstuck. If these builds are queued behind something that will never complete, please clear it.
Impact on us: the currently-running containers are healthy and still serving (build 97e13aae,
booted 2026-08-13), so our customers are unaffected. But we cannot ship any code change to either
service until builds work, and we have a performance fix and a safety fix waiting.
7 Replies
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
Your builds weren't stuck on a broken builder. Every one of them failed at the very first step, downloading your repo from GitHub, which kept refusing the request with an authorization error. That step retries quietly for up to an hour, which is why you only ever saw the one scheduling line, and why cancel and redeploy asked you to wait. The last build that was still running has failed out on its own now, and nothing is queued behind these.
Two things overlapped today. The morning failures match the authorization problem you hit on the trigger side and fixed by reinstalling the app. Separately, GitHub has been in a major incident since about 13:40 UTC with sporadic authentication failures, and we have an incident up for it at status.railway.com/incident/W4MIGEVT. We can't tell from our side whether your newest failures are that incident or a leftover of the install, so here's the test: once GitHub declares recovery, push or redeploy. If that build still fails, reply here and we'll dig into the app installation with you.
Railway Team
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
Thanks - we ran the test you asked for. GitHub has declared full recovery (status page all-operational, the Aug 17 auth incident resolved), and we then triggered one fresh deploy per service at 18:55 UTC today: d8d90925-6269-40cb-b536-bec15281ccd6 (titan-workers) and 468fe63c-6edb-49f6-b19a-1585b6bbf6f7 (titan-web-design). Over 20 minutes later both still sit at the single "scheduling build on Metal builder" line, same as every build since this began. So per your test, this is not the GitHub incident - please dig into the app installation with us.
Detail that may help:
- Our failures began 06:14 UTC Aug 17, BEFORE the GitHub incident window you referenced (~13:40 UTC), so the incident can't be the original cause.
- The GitHub App was fully uninstalled and reinstalled on Aug 17, and the repo shows the app authorized - yet the clone step still gets refused.
- Our side is verified end to end since the reinstall: both services carry the correct Dockerfile manifests and config-as-code paths, watch patterns behave (docs-only pushes correctly skip), buildEnvironment pinned to V2.
- A brand-new throwaway service on the same repo wedged identically on a third distinct builder, so it isn't per-service state.
What should we check in the installation, or can you inspect the grant on your side? Happy to reinstall again if you tell us what to verify afterward.
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
Update: running an investigation now as well, the behavior you are seeing definitely is not intentional.
Status changed to Awaiting User Response Railway • about 2 months ago
2 months ago
QQ: Off the top of your head, do you know how big is the build artifact you are uploading?
2 months ago
Yup, that seems to be it, sorry for the wrong path of investigation, you committed some giant build artifacts which is giving you some opaque errors. (Which we would make more clear.)
Our build step treats a 20 second gap in the incoming byte stream as a stalled job and restarts it, and the restart begins again from zero bytes, so a download that big never gets to the end. (Your big repo...)
Your builds ran that loop up to 7 times each, which was the wait.
Second, the GitHub credential we mint for the download lives exactly one hour and is not refreshed between those retries. Once the loop crosses the hour mark, GitHub rejects the now-expired credential, and the error we surface for that rejection is "repository not authorized". So the message you have been chasing is the last symptom of a size and timeout problem, not a permissions problem.
Can you run this
`
git ls-tree -r -l HEAD | sort -k4 -n -r | head -30
git ls-tree -r -l HEAD | awk '{s+=$4} END {print s/1e9, "GB total"}'`
And see what is causing the repo size to blow up?
2 months ago
Here is what those two commands returned BEFORE the fix:
git ls-tree -r -l HEAD | awk '{s+=$4} END {print s/1e9, "GB total"}'
13.7503 GB total
The top of the largest-files list was all committed screenshots, around 15 MB each:
15219953 scratchpad/hq49/pool-round2/live/pool-home-393-full.png
15017868 scratchpad/hq49/pool-round2/after/m-home-full.png
14722623 scratchpad/hq53/general-contractor/frames/home-390-full.png
14709992 scratchpad/hq53/general-contractor/frames/mobile-home-390-full.png
14432567 scratchpad/_hq53/auto-repair/frames/home-390-full.png
Broken down, the repo was 45,796 files and the weight was almost entirely images: 10.19 GB across 12,974 .png files, plus 0.53 GB across 821 .jpg files. By directory that was scratchpad/ at 13.02 GB, engine/scratchpad/ at 0.49 GB, and everything else at just 0.24 GB.
So 98% of our repo was automated QA screenshots. Our build agents capture full-page screenshots of every template at desktop and phone width as visual evidence, and those were being committed alongside the code. Not build artifacts in the usual sense, but the same effect on your download.
WHAT WE CHANGED: we untracked every image under those two directories (git rm --cached, so the files stay on disk and stay in history - nothing destroyed) and added them to .gitignore. That landed on main as commit 19d0904d.
git ls-tree -r -l HEAD | awk '{s+=$4} END {print s/1e9, "GB total"}'
0.455 GB total
HEAD is now 0.455 GB across 30,027 files, down from 13.75 GB - about a 30x reduction.
ONE QUESTION BACK, because it decides whether we are done:
Does your build step download a HEAD-only snapshot (a tarball for the commit, or a shallow --depth 1 clone), or does it do a full clone with history?
If it is HEAD-only, we should be fixed as of that commit and you can retry our build. But if you clone full history, the old image blobs are still in there - our packed object store is 1.33 GiB packed plus a large amount of loose objects not yet garbage-collected - and the download will still be far bigger than it should be. In that case we will run a history rewrite (git-filter-repo or BFG) to strip those blobs out of history and force-push, which we would rather schedule deliberately than discover on the next failed deploy.
Status changed to Awaiting Railway Response Railway • about 2 months ago
2 months ago
We pull only the specific commit for builds, not the full repo history. Your cleanup is the complete fix. You can deploy now.
Status changed to Awaiting User Response Railway • about 2 months ago
a month ago
This thread has been marked as solved automatically due to a lack of recent activity. Please re-open this thread or create a new one if you require further assistance. Thank you!
Status changed to Solved Railway • about 1 month ago