Cron job killed mid-execution — no exit code, no SIGTERM, low resource usage
lupinirudy
HOBBYOP

2 months ago

Subject: Cron job killed mid-execution — no exit code, no SIGTERM, low resource usage

Hi,

Our cron service (project: truthful-delight, service: backup-notturno,

region: europe-west4-drams3a) has been failing every night since Aug 15.

It ran green from creation until Aug 14.

What the job does: an rclone copy to encrypted remote storage, followed

by an rclone cryptcheck integrity verification. Both are standard rclone

commands in an alpine:3.20 container, ENTRYPOINT is a bash script.

The symptom:

  • The copy step always completes successfully (logged).
  • The verification step starts (logged: "cryptcheck in corso...").
  • Then the container dies. No further output at all.

Today we deployed a change that prints a summary from a bash EXIT trap,

specifically so that the exit code and diagnostics would survive any

failure. We verified locally that this trap fires on normal exit, on

set -e failures, and on SIGTERM (exit 143) — it only fails to fire on

SIGKILL.

After deploying, we triggered two manual runs. Neither produced the trap

output. So the container is being SIGKILLed, not terminated gracefully.

What rules it out:

  • Not OOM: Metrics show memory peaking under 100 MB, CPU near zero for

    the whole run.

  • Not a fixed timeout: run durations vary — 45s, 47s, 1m13s, and 17s.

  • Not our code: the script cannot SIGKILL itself, and the trap would

    catch any exit it controls.

Configuration: cronSchedule "0 1 * * *", numReplicas 1,

restartPolicyType NEVER, runtime V2, builder DOCKERFILE, root directory

/backup-service.

Questions:

  1. What is the actual container exit code for these runs? We cannot see

    it in the dashboard — only "Crashed".

  2. Is there a platform-level limit (execution time, network egress,

    process behaviour) that would SIGKILL a cron container under these

    conditions?

  3. Did anything change in the V2 cron runtime around Aug 14-15, 2026?

    That is exactly when the job started failing, with no change on our

    side (our last code change to this service was July 29).

Thanks.

$10 Bounty

3 Replies

Railway
BOT

2 months ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • about 2 months ago


lofimit
HOBBY

2 months ago

Hi,

Could you provide the termination reason/exit code reported by Railway for one of the failing runs? (The platform may give another reason that isnt SIGKILL)


lupinirudy
HOBBYOP

2 months ago

Hi, thanks for picking this up.

That's the problem: Railway doesn't show me a termination reason or an

exit code anywhere I can find. The deployment page shows only the badge

"Crashed" — no exit code, no termination reason. The Cron Runs list shows

the run with a red dot and its duration, nothing more. I've looked under

Details, Deploy Logs, and the deployment's overflow menu.

If there's a place in the dashboard (or an API/CLI call) that exposes the

termination reason, please point me to it and I'll report it here — that's

exactly the number I'm missing.

For reference, the deploy logs for a failing run contain three lines

total: container start, "copy completed", "verification starting". Then

nothing.


lupinirudy
HOBBYOP

2 months ago

Solved — and it wasn't the platform. The summary I mentioned deploying

did fire; I read the logs of a run that was still executing the previous

code, so its absence proved nothing. With the new code the run prints:

"cryptcheck FAILED: differences found: 1 (rclone exit: 1). Files checked:

145" plus the summary and "script exit: 1". So the container exits

normally with code 1 — rclone cryptcheck is reporting one real difference

between source and encrypted remote. Thanks for the nudge toward the

termination reason: chasing it is what made me re-read the logs properly.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...