TT Lab
Get started
Learn Learning paths Courses

Linux Incident Response

Whom Does the OOM Killer Choose

Continue in TT Lab

In one line

The OOM Killer's score is based on RSS, not on top's VIRT, and oom_score_adj is an adjustment whose unit is one thousandth of total memory. And exit code 137 only means "it died from SIGKILL"; it does not mean OOM.

Why this was needed

"The service died but there is no error in the application log. The last line is a normal request being processed."

In this kind of report, the log being clean is itself a clue. SIGKILL cannot be caught and no handler runs, so the application has no chance to leave anything. The trace is left not by the application but by the kernel.

There are four common misunderstandings here.

  1. 137 does not mean OOM. It only means 128+9, that is, it died from SIGKILL. Expiry of the grace period after a liveness probe failure, a manual kill and a node shutdown are all 137 too.
  2. The culprit is not VIRT. A process that merely reserves 64GB with mmap and actually uses only 200MB is not even a candidate.
  3. "Adding swap prevents OOM" - only half true. It doesn't prevent it, it delays it, and the thrashing in that interval can be worse than dying instantly.
  4. A container can be OOMKilled without exceeding its limit. The same indication appears when a node-wide OOM kills a process inside the container. In this case, raising the limit won't stop it from recurring.

How it works

The score calculation is roughly this.

points = rss + swapents + (pgtables / PAGE_SIZE)      # 단위: 페이지
adj    = oom_score_adj * (totalpages / 1000)
badness = points + adj

That is, an oom_score_adj of 200 is a bonus worth about 12.8GiB on a 64GiB machine. A 2GB RSS process can beat an 11GB JVM with this bonus. Conversely, -1000 is special - it is excluded from the score calculation entirely and is never chosen. That is why distributions give this value to sshd.

When does it trigger? Not "when memory reaches 0" but "when reclaim was attempted and made no progress." That is why, if swap exists, reclaim makes progress and OOM is postponed, and thrashing arises in that interval.

A global OOM and a cgroup OOM are different events. In the oom-kill: line at the end of the dmesg report, the constraint tells them apart. CONSTRAINT_NONE is global and CONSTRAINT_MEMCG is cgroup. Only for a cgroup OOM does the memory: usage/limit/failcnt line appear with it. This one line completely separates the responses - if global, you have to look at the placement of the whole node, and if cgroup, it is a matter of that container's limit and actual usage.

A container's real limit is not free. Inside a container, free -m shows the whole host. The actual ceiling is /sys/fs/cgroup/memory.max and the current usage is memory.current. And the three counters in memory.events tell you the state - max is the number of times it hit the ceiling and reclaimed, oom is the number of times it entered the OOM path after reclaim failed, and oom_kill is the number of actual kills. If only max is large and oom_kill is 0, it is still holding on, so catching this is better than a post-mortem.

What you see in the field

MemFree and MemAvailable are different. Linux uses spare memory as page cache. A small MemFree is normal, and the criterion for judgment is MemAvailable, which counts in the reclaimable cache. In the free command as well, you look not at the free column but at the available column.

Adding up RSS gives more than the real value. RSS counts the library pages shared by several processes in full for each one. Simply adding up for a web server with 20 workers gives several times the real figure. For capacity estimation, use PSS (/proc/PID/smaps_rollup), which divides shared pages by the number of processes referencing them.

A Kubernetes request is not a number for fooling the scheduler but a survival ranking. The oom_score_adj of a Burstable Pod is calculated from the request ratio, so a Pod that declares a small request and uses a lot is the first to die under node pressure.

How to find out later why it died

The OOM killer is not silent. You just have to know where to look.

dmesg -T | grep -iE 'killed process|out of memory|oom-kill'
journalctl -k --since "1 hour ago" | grep -i oom

The kernel message leaves not only the killed process but also the list of candidates at that moment. If you look at the total-vm, anon-rss and oom_score_adj columns, who was using how much comes out as it is.

In containers the cgroup kills first. This is where the host as a whole has room but only that Pod dies. It appears together with exit code 137 (128+9).

cat /sys/fs/cgroup/<경로>/memory.events   # oom_kill 이 0이 아니면 이미 죽었다
kubectl get pod <파드> -o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'

Sometimes it is OOMKilled but the container seems not to have died. If only a child process inside the Pod dies, PID 1 is still alive, so there is no restart. The service is then in a state of "alive but processing nothing". What catches this is the readiness probe.

Cache is not memory in use. In free -h, judge by available. If you look only at used, you misread the normal state of a full cache as a shortage. Conversely, if available is small but buff/cache is large, that cache may be memory that can't be reclaimed (tmpfs, mlock, dirty pages).

Sometimes the problem is not memory but reservation. Depending on the overcommit setting, malloc succeeds and the process dies later when it actually uses the memory. That is how the phenomenon "allocated, but dies later" arises.

cat /proc/meminfo | grep -E 'MemAvailable|Committed_AS|CommitLimit'
sysctl vm.overcommit_memory vm.overcommit_ratio

You can adjust priority. For a process that must never die, lower oom_score_adj, and for one that may die first, raise it. But the root fix is setting limits accurately, and adjusting the score is a stopgap in the meantime.

What you will do in the next lab

You read MemTotal and MemAvailable directly from /proc/meminfo and confirm the difference between them. You start a process that really holds memory and see how VmSize and VmRSS differ, and raise oom_score_adj to observe the score actually change. At the end you build a tool that picks the top RSS processes and a script that watches for threshold breaches.