Railway Node.js Deployment Killed with Exit Code 137 — Possible OOM?
hovigvs
FREEOP

2 days ago

Hi Railway Community,

I'm troubleshooting a Node.js 24 application deployed on Railway that unexpectedly terminates during startup.

The application uses PGlite 0.3.16 with a persistent Railway volume.

What happened:

The application completed its initial database initialization and schema checks.

During the next initialization phase, the process was unexpectedly killed.

Deployment logs show:

Killed

[ELIFECYCLE] Command failed with exit code 137.

The process ran for approximately two seconds before termination.

Railway's CPU and memory metrics do not show a significant spike, possibly because the process terminated too quickly for the graphs to capture it.

Questions:

Is exit code 137 commonly associated with OOM kills on Railway?

Is there a way to retrieve historical OOM-kill information or container termination reasons?

How can I verify the effective memory and CPU limits for a deployment?

Could Railway terminate a process with SIGKILL for reasons unrelated to memory exhaustion?

I'm trying to diagnose the issue before modifying resources or redeploying.

Any guidance would be appreciated. Thank you!

$10 Bounty

3 Replies

Railway
BOT

2 days ago

This thread has been opened as a bounty so the community can help solve it.

Status changed to Open Railway • 2 days ago


2 days ago

Yea, the 137 code is an OOM kill signature.

As for the graph not showing it, you'll want to check the metrics tab per replica if you have multiple replicas too

Otherwise, yes, it's very possible Railway's resource telemetry for the graphs simply missed the spike.

You'll want to find anything that could spike your memory.

Also worth increasing --max-old-space-size for your node process.


3rkes
FREE

5 hours ago

137 with “Killed” is consistent with SIGKILL (128 + signal 9). OOM is one possible source of that signal; the exit code alone cannot prove it. A supervisor or manual kill can produce the same result. I haven't inspected your deployment, so I can't identify which happened here.

For a read-only check in an already-running container, if its cgroup v2 files are exposed at /sys/fs/cgroup, inspect:

for f in memory.max memory.current memory.peak memory.events memory.events.local cpu.max; do
  if [ -r "/sys/fs/cgroup/$f" ]; then
    printf '\n%s\n' "$f"
    cat "/sys/fs/cgroup/$f"
  fi
done

memory.max is the local hard limit in bytes (max means no limit at that level); a parent cgroup may impose a lower effective limit. cpu.max gives quota and period; quota/period is the CPU allowance, or max if uncapped locally. Missing files require checking the cgroup mount/layout rather than concluding there is no limit.

An increase in oom_kill in memory.events around the failure is stronger OOM evidence than the graph. The counters are cumulative, not timestamped; an existing nonzero value does not identify this particular crash. memory.peak, when available, can catch a brief spike. These observations only describe that cgroup: a replacement container cannot reconstruct a destroyed deployment's counters. For the historical event, ask Railway to correlate the exact deployment ID and UTC failure time with its termination records.

Before increasing --max-old-space-size, confirm the container budget. That flag limits V8 old space, not the whole process; increasing it can worsen a container OOM. In a later controlled diagnostic run, log process.memoryUsage() around each initialization step to distinguish heap growth from RSS/external allocations. That requires a code change, so it is separate from your requested read-only investigation. Avoid changing limits or redeploying merely to obtain evidence of the old crash.

Sources: Node exit codes, Node heap flag, Linux cgroup v2, Railway scaling.

Disclosure: prepared with AI assistance and checked against the linked official documentation; not tested in your Railway deployment.


hovigvs
FREEOP

4 hours ago

Thank you for the clarification.

The affected deployment is currently stopped, and I do not want to restart it simply to collect diagnostic information.

Could Railway confirm from its historical deployment or infrastructure records:

Whether the process was terminated by a container OOM event, platform intervention, or another SIGKILL source.

The effective memory limit in bytes at the time of failure.

Whether any memory peak or OOM event information is available for the failed deployment.

The Node.js runtime version and container architecture, if available.

I can provide the complete deployment ID and UTC failure timestamp privately if needed.

I'm trying to establish the cause without changing production settings or triggering another deployment.

Thank you.


Welcome!

Sign in to your Railway account to join the conversation.

Loading...