2 months ago
Subject: Cron job killed mid-execution — no exit code, no SIGTERM, low resource usage
Hi,
Our cron service (project: truthful-delight, service: backup-notturno,
region: europe-west4-drams3a) has been failing every night since Aug 15.
It ran green from creation until Aug 14.
What the job does: an rclone copy to encrypted remote storage, followed
by an rclone cryptcheck integrity verification. Both are standard rclone
commands in an alpine:3.20 container, ENTRYPOINT is a bash script.
The symptom:
- The copy step always completes successfully (logged).
- The verification step starts (logged: "cryptcheck in corso...").
- Then the container dies. No further output at all.
Today we deployed a change that prints a summary from a bash EXIT trap,
specifically so that the exit code and diagnostics would survive any
failure. We verified locally that this trap fires on normal exit, on
set -e failures, and on SIGTERM (exit 143) — it only fails to fire on
SIGKILL.
After deploying, we triggered two manual runs. Neither produced the trap
output. So the container is being SIGKILLed, not terminated gracefully.
What rules it out:
-
Not OOM: Metrics show memory peaking under 100 MB, CPU near zero for
the whole run.
-
Not a fixed timeout: run durations vary — 45s, 47s, 1m13s, and 17s.
-
Not our code: the script cannot SIGKILL itself, and the trap would
catch any exit it controls.
Configuration: cronSchedule "0 1 * * *", numReplicas 1,
restartPolicyType NEVER, runtime V2, builder DOCKERFILE, root directory
/backup-service.
Questions:
-
What is the actual container exit code for these runs? We cannot see
it in the dashboard — only "Crashed".
-
Is there a platform-level limit (execution time, network egress,
process behaviour) that would SIGKILL a cron container under these
conditions?
-
Did anything change in the V2 cron runtime around Aug 14-15, 2026?
That is exactly when the job started failing, with no change on our
side (our last code change to this service was July 29).
Thanks.
3 Replies
2 months ago
This thread has been opened as a bounty so the community can help solve it.
Status changed to Open Railway • about 2 months ago
2 months ago
Hi,
Could you provide the termination reason/exit code reported by Railway for one of the failing runs? (The platform may give another reason that isnt SIGKILL)
2 months ago
Hi, thanks for picking this up.
That's the problem: Railway doesn't show me a termination reason or an
exit code anywhere I can find. The deployment page shows only the badge
"Crashed" — no exit code, no termination reason. The Cron Runs list shows
the run with a red dot and its duration, nothing more. I've looked under
Details, Deploy Logs, and the deployment's overflow menu.
If there's a place in the dashboard (or an API/CLI call) that exposes the
termination reason, please point me to it and I'll report it here — that's
exactly the number I'm missing.
For reference, the deploy logs for a failing run contain three lines
total: container start, "copy completed", "verification starting". Then
nothing.
2 months ago
Solved — and it wasn't the platform. The summary I mentioned deploying
did fire; I read the logs of a run that was still executing the previous
code, so its absence proved nothing. With the new code the run prints:
"cryptcheck FAILED: differences found: 1 (rclone exit: 1). Files checked:
145" plus the summary and "script exit: 1". So the container exits
normally with code 1 — rclone cryptcheck is reporting one real difference
between source and encrypted remote. Thanks for the nudge toward the
termination reason: chasing it is what made me re-read the logs properly.
