TT Lab
Get started
Learn Learning paths Courses

Linux Incident Response

Diagnosing Load and CPU

Continue in TT Lab

Goal

You confirm for yourself what the load average counts, and by reading CPU utilization, PSI and the throttling counter together, you become able to explain a situation where "the metrics are normal but it is slow".

Why it matters

Linux's load average does not count only CPU waiting. It also counts tasks in the D state that are waiting for disk or NFS responses. So an 8-core server with a load of 24 and a CPU idle of 60% is normal behavior. Adding CPU then just spends money. Conversely, a container's cgroup throttling appears neither in load, nor in steal, nor in utilization - because a throttled task is removed from the run queue entirely. That is why you have to look separately at the counters in cpu.stat. The reason this lab even grades "if it isn't there, write that it isn't" is that in the field too, the first action is to check whether that file exists.

Steps

Work in the /root/load directory.

  1. Write the number of CPUs in this environment, as a number only, to /root/load/cpus.txt.
  2. Start 2 or more processes that keep using the CPU in the background and write their PIDs to /root/load/burner_pids.txt, one per line. They must stay alive through all the following steps.
  3. After waiting 1–2 minutes, write the 1-minute average from /proc/loadavg to /root/load/load1.txt. It must be 0.7 or more.
  4. Write one letter for the current state of those processes to /root/load/state.txt.
  5. Write the cumulative CPU time (seconds) of the first process to /root/load/cpu_sec.txt as an integer. It must be 5 seconds or more.
  6. In /proc/pressure/cpu, on the some line, take the avg10 value and write it to /root/load/psi.txt. If the file doesn't exist, write unavailable.
  7. In the cgroup's cpu.stat, take the nr_throttled value and write it to /root/load/throttle.txt in the form nr_throttled=<값>. If it can't be read, write unavailable.
  8. Create /root/load/loadcheck.sh <임계비율> (the argument is the threshold ratio). If 1분 부하 평균 / CPU 개수 (the 1-minute load average divided by the CPU count) exceeds the threshold ratio, print high load=<값> cpus=<개수> and exit with code 1; if it doesn't, print ok and exit with code 0. With no argument, exit with a non-zero code.

Notes

Check the number of CPUs

Write the number of CPUs in this environment, as a number only, to /root/load/cpus.txt.

nproc is the simplest. The load average only has meaning when placed next to this number.

Start processes that burn CPU

Start 2 or more processes that keep using the CPU in the background and write their PIDs to /root/load/burner_pids.txt, one per line. They must stay alive through all the following steps.

Two python3 processes running an infinite loop are enough. Write the PIDs one per line.

Confirm that the load average rises

After waiting 1–2 minutes, write the 1-minute average from /proc/loadavg to /root/load/load1.txt. It must be 0.7 or more.

It is an exponentially decaying average, so it doesn't rise immediately. Wait 1–2 minutes and then read the first value of /proc/loadavg.

Read the process state letter

Write one letter for the current state of those processes to /root/load/state.txt.

The third field of /proc//stat is the state. What letter would a process that keeps using the CPU have?

Measure cumulative CPU time

Write the cumulative CPU time (seconds) of the first process to /root/load/cpu_sec.txt as an integer. It must be 5 seconds or more.

Add utime and stime from /proc//stat and divide by the clock tick (usually 100).

Read the PSI pressure metric

In /proc/pressure/cpu, on the some line, take the avg10 value and write it to /root/load/psi.txt. If the file doesn't exist, write unavailable.

The some line of /proc/pressure/cpu has avg10. If the file doesn't exist, write unavailable.

cgroup throttling counter

In the cgroup's cpu.stat, take the nr_throttled value and write it to /root/load/throttle.txt in the form nr_throttled=<값>. If it can't be read, write unavailable.

cpu.stat has nr_periods and nr_throttled. If it can't be read, it is unavailable.

Load judgment script

Create /root/load/loadcheck.sh <임계비율> (the argument is the threshold ratio). If 1분 부하 평균 / CPU 개수 (the 1-minute load average divided by the CPU count) exceeds the threshold ratio, print high load=<값> cpus=<개수> and exit with code 1; if it doesn't, print ok and exit with code 0. With no argument, exit with a non-zero code.

Compare the ratio of the load average divided by the core count with the threshold. You can compare decimals with awk even without bc.