TT Lab
Get started
Learn Learning paths Courses

Linux Incident Response

Load Average Is Not CPU Utilisation

Continue in TT Lab

In one line

Linux's load average does not count only CPU waiting. It also counts tasks in the D state that are waiting for disk or NFS responses. So load can be 24 even when the CPU is idle, and then the culprit is storage.

Why this was needed

An alarm at dawn. The load average of an 8-core server is 24.31, 22.08, 14.95, yet the CPU idle in top is over 60%. Here it splits two ways - "monitoring is wrong" or "we need 3 times the CPU". Both are wrong.

Even the units differ. Utilization is a ratio (0–100%), while load is a count (number of tasks). And "load = CPU run queue length" is true on other Unixes but false on Linux.

How it works

The kernel samples each CPU's run queue once every 5 seconds.

nr_active = 실행 중이거나 실행 대기(R) + 중단 불가 대기(D)

In 1993, Linux added TASK_UNINTERRUPTIBLE (the D state) to this calculation. The intent was to measure demand for system resources overall, not only CPU demand. The D state is a kernel-internal wait that can't even be woken by a signal - block I/O completion, page cache locks, hard-mounted NFS response waits, some kernel mutexes. The CPU is not on this list.

And this value is not an arithmetic mean but an exponentially decaying average. If load rises like a step, the 1-minute average reflects only about 63% of it after 1 minute, and it takes about 5 minutes to reach 99%. Two things follow from this.

The diagnostic order is like this.

cat /proc/loadavg          # 네 번째 필드 '실행가능/전체'를 함께 본다
ps -eo state --no-headers | sort | uniq -c   # R 과 D 를 센다
vmstat 1 5                 # b 열이 D 상태 개수, wa 가 I/O 대기 비율
iostat -x 1 3              # await 와 큐 깊이로 판단 (%util 100% 는 포화가 아닐 수 있다)

If load is 24 and R is 2, the remaining 22 are D. Then the answer lies not in the CPU but on the storage side.

You also need to know where "divide by the core count" breaks down.

  1. D-state contamination - exactly as seen above
  2. Lag of the average - a 30-second spike doesn't show even half in the 1-minute average
  3. SMT - an nproc of 16 may be 8 physical cores
  4. cgroup throttling - a throttled task is removed from the run queue entirely, so it is not captured in load
  5. No attribution information - it is a single global scalar, so all an alarm can do is tell you to "go in and look"

What you see in the field

The silent stall of a container. If a cgroup has a CPU cap, the kernel distributes a quota every period (default 100ms), and if the group exhausts it within the period, the whole group is pulled off the run queue for the remaining time. With a limit of 1 core and 8 threads, the 100ms quota is exhausted in 12.5ms and the remaining 87.5ms it stands completely still. Yet the average utilization looks like about 12.5%, and it isn't caught in load or in steal. Only the cpu.stat ratio nr_throttled / nr_periods tells you this. Above 5% is worth investigating, and above 10% calls for action.

Steal time. In top, st is "time I wanted to use but the hypervisor gave to another guest". No matter how much you optimize the application, this number won't go down. Below 1% is normal, 5–10% sustained calls for action, and above 10% no tuning on that host is effective. And running reboot inside the guest doesn't change the physical host - only a stop/start through the cloud API relocates it.

The question PSI answers. Load says only "how many are waiting" and not "how long they were stalled". In /proc/pressure/{cpu,io,memory}, some is the fraction of time at least one task was stalled, and full is the fraction of time all of them were stalled at once. An io full avg10=61 means that in the last 10 seconds, the whole system was stopped because of I/O for more than 6 seconds, and this translates directly into lost throughput.

When load is high but the CPU is idle

A situation where load average is 40 but the CPU utilization in top is 5% is not a malfunction. Linux load counts not only waiting to run but also the D state (waiting on disk). So high load means not "busy" but "there is a lot waiting for something".

First sort out which kind it is. In vmstat, the r column is waiting to run and the b column is blocked waiting.

vmstat 1 5
# r  b   swpd  free  buff cache  si so  bi bo  in cs  us sy id wa st

If r is large, the CPU is short, and if b and wa are large, storage is the bottleneck. If st (steal) is not 0, the virtual machine is not getting the physical CPU, so it is a problem on the host side.

When the CPU is short, compare with the core count. A load of 4 is serious on 1 core and idle on 8 cores. The criterion is whether the value divided by nproc exceeds 1. However, nproc inside a container often reports the host's core count, so also look at the cgroup's cpu.max.

Check whether one process is using a whole core or several threads are sharing it.

top -H -p <pid>          # 스레드 단위
pidstat -t -p <pid> 1

If the CPU utilization is stuck at exactly 100%, it is a single thread using a whole core, and then adding cores does nothing.

Don't miss throttling. If a CPU limit is applied in a container, utilization looks low but responses are slow. The cgroup statistics record how many times it happened.

cat /sys/fs/cgroup/cpu.stat     # nr_throttled, throttled_usec

If nr_throttled keeps increasing, raise the limit or change how requests are handled to reduce short bursts.

iowait only means "the CPU is idle and meanwhile waiting on the disk". If the CPU is busy, wa comes out low even with the same disk latency. Whether storage is slow is judged, in iostat -x 1, by await and %util.

What you will do in the next lab

You start a process that burns CPU and watch the load average actually rise, and read the R state and the cumulative CPU time directly from /proc. You check PSI and the cgroup throttling counter, and if the environment doesn't have them, you record precisely that they are absent. At the end you build a script that normalizes load by the core count to make a judgment.