Blog
6 min read

The Linux OOM Killer: Why Your Process Was Killed

When Linux runs out of memory it picks a process and kills it. How the kernel OOM killer chooses its victim, oom_score and oom_score_adj, cgroup v2 memory limits versus system-wide OOM, overcommit, reading the dmesg report, exit code 137 in containers, and how to stop it happening.

Your app was running, and then it wasn't. No stack trace, no error log — the process just vanished, or your container restarted with exit code 137. Nine times out of ten, the Linux OOM killer did it.

When the kernel cannot find memory to satisfy an allocation and can't reclaim enough (by dropping page cache, swapping, or compacting), it has two choices: panic, or kill something. By default it kills something. (Docker container keeps restarting)

Why memory runs out "suddenly": overcommit

Linux hands out virtual memory optimistically. malloc of 1 GB usually succeeds immediately even if you don't have 1 GB free, because pages are only backed by physical RAM when first touched. This is overcommit, controlled by vm.overcommit_memory:

Value Behaviour
0 (default) Heuristic — refuse only obviously absurd allocations
1 Always allow; never refuse
2 Strict — commit limit is swap + overcommit_ratio% of RAM; allocations beyond it fail with ENOMEM

Because allocations rarely fail up front, the failure moves to the moment the memory is used — a page fault deep inside your program, where there's no error to return. The kernel's only remedy at that point is to free memory by killing a process.

Mode 2 turns OOM kills into allocation failures, which sounds nicer, but most software (including runtimes like Node and the JVM, which reserve large virtual ranges) handles ENOMEM badly. It's mostly used on dedicated database hosts.

How the victim is chosen

The kernel computes an oom_score for each process, roughly proportional to how much memory it would free: resident memory plus swap plus page tables, as a fraction of what's available. The highest score dies.

You can see and influence it:

cat /proc/<pid>/oom_score        # current score, 0–1000 (ish)
cat /proc/<pid>/oom_score_adj    # your adjustment, -1000 to 1000
echo -1000 | sudo tee /proc/<pid>/oom_score_adj   # never kill this process
echo 500 | sudo tee /proc/<pid>/oom_score_adj     # prefer killing this one
  • -1000 exempts a process entirely. Use it sparingly — sshd is a reasonable candidate, a leaky app server is not.
  • systemd units can set OOMScoreAdjust= instead of poking /proc.
  • Children inherit the parent's adjustment.

The old heuristics (bonuses for root, penalties for long-running processes) are long gone; the modern algorithm is close to "biggest memory user, adjusted".

Reading the kernel's report

Every OOM kill writes a detailed report to the kernel log:

sudo dmesg -T | grep -iE 'out of memory|oom-kill|killed process'
journalctl -k | grep -i oom

Typical lines:

node invoked oom-killer: gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0, oom_score_adj=0
oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=...,mems_allowed=0,oom_memcg=/system.slice/app.service,...
Memory cgroup out of memory: Killed process 4121 (node) total-vm:2843120kB, anon-rss:1012344kB, file-rss:0kB, shmem-rss:0kB, UID:1000 pgtables:3120kB oom_score_adj:0

Things to read:

  • The process that invoked the killer isn't necessarily the victim — it's just whoever tried to allocate when memory ran out.
  • constraint= tells you the scope. CONSTRAINT_NONE means the whole machine ran out. CONSTRAINT_MEMCG means a cgroup hit its limit — the machine may have had plenty free.
  • anon-rss of the victim is the memory it actually held. If it's huge, you have your culprit. If it's small and the host still ran out, something else is eating memory (often many small processes, or tmpfs).
  • The report also dumps a table of every process with its RSS — invaluable when the victim isn't the real hog.

cgroup OOM vs system OOM

On modern systems most OOM kills are cgroup kills, not global ones. Docker's --memory, Kubernetes resources.limits.memory, and systemd's MemoryMax= all set a cgroup v2 memory.max. When processes in that group exceed it and reclaim fails, the kernel OOM-kills inside that group only. (Container CPU and memory limits with cgroups)

Useful files in the cgroup directory:

cat /sys/fs/cgroup/<group>/memory.current   # usage now
cat /sys/fs/cgroup/<group>/memory.max       # hard limit
cat /sys/fs/cgroup/<group>/memory.events    # oom, oom_kill counters

cgroup v2 also has memory.high — a soft limit where the kernel throttles and reclaims aggressively before the hard limit. Setting memory.high a bit below memory.max turns sudden kills into slowdowns you can see coming. And memory.oom.group = 1 kills the whole group together rather than one process, which is what you want when a half-dead set of workers is worse than a clean restart.

Exit code 137 in containers

A process killed by SIGKILL exits with 128 + 9 = 137. In Docker:

docker inspect <container> --format '{{.State.OOMKilled}} {{.State.ExitCode}}'

OOMKilled: true confirms it. In Kubernetes, the pod status shows Reason: OOMKilled. Note that 137 alone doesn't prove OOM — anything that sends SIGKILL (a docker kill, a failed graceful-shutdown timeout) gives the same code. (Graceful shutdown in Node.js)

The runtime trap: heap limits vs container limits

Language runtimes decide their own heap size, and they don't always respect the container:

  • Node.js sizes its default old-space heap from available memory; recent versions are cgroup-aware, but if you set --max-old-space-size higher than the container limit, the kernel kills you before V8 ever throws "heap out of memory". Keep it at roughly 75% of the limit to leave room for buffers, native modules and the stack. (JavaScript heap out of memory)
  • JVM — use -XX:MaxRAMPercentage rather than a fixed -Xmx.
  • Python has no heap cap at all; a leak grows until the cgroup stops it.

If you see OOM kills with no runtime error first, the runtime thought it had more room than the kernel did. (Node memory leaks)

systemd-oomd and earlyoom

The kernel OOM killer acts late — only after reclaim has failed, which often means minutes of the system thrashing first. Userspace daemons act earlier:

  • systemd-oomd (default on Fedora and Ubuntu desktop) watches memory pressure (PSI) and swap usage per cgroup and kills the whole offending cgroup when pressure stays high. (Linux pressure stall information)
  • earlyoom is a simpler daemon that kills the largest process when free memory and swap drop below thresholds.

On a server these trade a frozen machine for a faster, more predictable kill.

Stopping it from happening

  1. Find the real consumer. Read the dmesg table, check memory.current per cgroup, and graph memory over time — a slow climb is a leak, a spike is a big request or batch job.
  2. Set limits deliberately. Every container should have a memory limit, so one runaway service is killed instead of taking the whole host (and your database) down with it.
  3. Protect what matters. Give your database and sshd a negative oom_score_adj; give batch jobs a positive one.
  4. Use memory.high to get throttling and alerts before kills.
  5. Add a little swap (or zram) on small VMs. It isn't a fix for too little RAM, but it absorbs short spikes and gives the kernel somewhere to put cold pages.
  6. Alert on oom_kill counters, not just on the app being down.

Every EasySpawn server is its own VM with dedicated memory, so another customer's leak can't trigger your OOM kill — and Claude Code on the server can dig through memory usage and container exit codes with you when something does die. See pricing or join the waitlist.

Related: Container CPU and Memory Limits · Docker Container Keeps Restarting · JavaScript Heap Out of Memory · Linux Pressure Stall Information

Keep reading