TT Lab
Get started
Learn Learning paths Courses

Diagnosing CPU and Memory Leaks

Load Average Is Not CPU Utilisation

Continue in TT Lab

In one line

The Linux load average is the number of runnable processes plus processes in uninterruptible sleep (D state). It rises when the disk is slow, too. It measures something different from CPU utilization.

Why this matters

"Load is 20, so let's add CPU" is a common misdiagnosis. On an 8-core machine, a load of 20 can mean the CPU is short, or it can mean NFS has stalled and 20 processes are stuck waiting on I/O. Adding CPU in the second case changes nothing.

Other Unix systems count only CPU waiters in load. Only Linux also counts D state (uninterruptible sleep). This change was made in 1993 to better reflect "the system is busy", and since then Linux load has been a measure of system demand, not a CPU metric.

How it works

The three numbers are exponential moving averages over 1, 5, and 15 minutes. The direction is the information.

cat /proc/loadavg
# 8.24 4.10 2.05 3/512 12345
#  └1분 └5분 └15분 └실행중/전체 스레드 └마지막 PID

If the 1-minute value is greater than the 15-minute value, things are getting worse right now; if it is the other way around, the system is recovering.

To see what is raising load, count processes by state.

ps -eo state,comm --no-headers | awk '{print $1}' | sort | uniq -c | sort -rn
# R = 실행 가능, D = 끊기지 않는 대기(대개 I/O), S = 잠듦

If there are many D processes, it is not a CPU problem.

To look at the CPU itself, read /proc/stat twice and take the difference.

awk '/^cpu /{print $2+$3+$4+$6+$7+$8, $5}' /proc/stat   # (busy, idle)

Utilization comes from the difference between the two readings. This is exactly what top does.

Common misconceptions

It is not a bug when top shows %CPU above 100. In the default mode, %CPU is relative to a single core. If 4 threads each run flat out, that is 400%. Pressing Shift+I in top switches to a percentage of all cores (turns off Irix mode).

top inside a container sees the whole host. Because /proc belongs to the host, both the core count and the load are host values. The share actually allocated to the container has to be read from the cgroup.

cat /sys/fs/cgroup/cpu.max        # "쿼터 주기" — max 면 제한 없음
cat /sys/fs/cgroup/cpu.stat       # nr_throttled, throttled_usec

If nr_throttled keeps growing, the CPU is not short; you are hitting a limit. The two call for completely different responses.

Load average is not a CPU metric

The Linux load average is the number of processes that are running or waiting to run, and it also includes processes in D state (uninterruptible sleep, usually disk I/O). This is where Linux differs from other Unix systems, and it is where the confusion comes from.

$ uptime
 load average: 8.42, 6.10, 4.33     ← 코어가 4개인데 8?
$ vmstat 1 3
 r  b   ...     ← r 은 실행 대기, b 는 I/O 대기
 1  7   ...     ← 실행 대기는 1뿐, 나머지 7은 디스크를 기다린다

If r is small and b is large, adding CPU improves nothing. The disk is the culprit. Scaling up an instance based on the load average alone is the most expensive misreading of this metric.

Same utilization, different causes

Which column is large in one line of top decides the direction of the investigation.

Large column Meaning Where to look first
us (user) Application code Profiler, hot functions
sy (system) Kernel — many system calls strace -c, file and socket usage patterns
wa (iowait) Waiting on the disk iostat -x, the await column
si (softirq) Network interrupts /proc/softirqs, RSS/RPS settings
st (steal) The hypervisor took it away Cloud instance class, noisy neighbors

If st exceeds 5%, it is not your problem. Either a burstable instance (the t family) has run out of credits, or the physical host is oversubscribed. No amount of code fixing will help.

The order for finding what uses the CPU

  1. Who — Drill down to threads with top -H -p <pid>. Often it is not the whole process but a single thread that is spinning (a GC thread, a polling loop).
  2. Where — Look at the hot symbols with perf top -p <pid>. If the symbols are full of [unknown], debug symbols are missing, so use a language-specific profiler.
  3. Why — Why that function is called so often is something you find in the code. Usually it is O(n²), the cache is not working, or logs are being written synchronously.

In a container without perf, /proc/<pid>/stack or the language runtime's thread dump (jstack, py-spy dump) can tell you half of it.

What really matters in practice

In Kubernetes, setting a CPU limit applies a cgroup quota. When requests pile up, throttling occurs and p99 latency spikes, but the CPU utilization graph looks flat near the limit, so it is easy to misread it as "there is headroom". Check throttling with cpu.stat, not with utilization.