Diagnosing Load and CPU
Goal
You confirm for yourself what the load average counts, and by reading CPU utilization, PSI and the throttling counter together, you become able to explain a situation where "the metrics are normal but it is slow".
Why it matters
Linux's load average does not count only CPU waiting. It also counts tasks in the D state that are waiting for disk or NFS responses. So an 8-core server with a load of 24 and a CPU idle of 60% is normal behavior. Adding CPU then just spends money. Conversely, a container's cgroup throttling appears neither in load, nor in steal, nor in utilization - because a throttled task is removed from the run queue entirely. That is why you have to look separately at the counters in cpu.stat. The reason this lab even grades "if it isn't there, write that it isn't" is that in the field too, the first action is to check whether that file exists.
Steps
Work in the /root/load directory.
- Write the number of CPUs in this environment, as a number only, to
/root/load/cpus.txt. - Start 2 or more processes that keep using the CPU in the background and write their PIDs to
/root/load/burner_pids.txt, one per line. They must stay alive through all the following steps. - After waiting 1–2 minutes, write the 1-minute average from
/proc/loadavgto/root/load/load1.txt. It must be 0.7 or more. - Write one letter for the current state of those processes to
/root/load/state.txt. - Write the cumulative CPU time (seconds) of the first process to
/root/load/cpu_sec.txtas an integer. It must be 5 seconds or more. - In
/proc/pressure/cpu, on thesomeline, take theavg10value and write it to/root/load/psi.txt. If the file doesn't exist, writeunavailable. - In the cgroup's
cpu.stat, take thenr_throttledvalue and write it to/root/load/throttle.txtin the formnr_throttled=<값>. If it can't be read, writeunavailable. - Create
/root/load/loadcheck.sh <임계비율>(the argument is the threshold ratio). If1분 부하 평균 / CPU 개수(the 1-minute load average divided by the CPU count) exceeds the threshold ratio, printhigh load=<값> cpus=<개수>and exit with code 1; if it doesn't, printokand exit with code 0. With no argument, exit with a non-zero code.
Notes
- You can burn CPU with
python3 -c "while True: pass" &. - The fields of
/proc/<PID>/statare in the orderpid (comm) state ..., and the 14th and 15th are utime/stime (clock ticks). - In the form
awk 'BEGIN {exit (a/b > t) ? 1 : 0}'you can compare decimals without bc. - Common mistake 1: if you read it right after creating the load in step 3, the value is low. An exponentially decaying average reflects only 63% even after 1 minute.
- Common mistake 2: if you clean up the load-generating processes first in step 8, the judgment test fails.
Check the number of CPUs
Write the number of CPUs in this environment, as a number only, to /root/load/cpus.txt.
nproc is the simplest. The load average only has meaning when placed next to this number.
Start processes that burn CPU
Start 2 or more processes that keep using the CPU in the background and write their PIDs to /root/load/burner_pids.txt, one per line. They must stay alive through all the following steps.
Two python3 processes running an infinite loop are enough. Write the PIDs one per line.
Confirm that the load average rises
After waiting 1–2 minutes, write the 1-minute average from /proc/loadavg to /root/load/load1.txt. It must be 0.7 or more.
It is an exponentially decaying average, so it doesn't rise immediately. Wait 1–2 minutes and then read the first value of /proc/loadavg.
Read the process state letter
Write one letter for the current state of those processes to /root/load/state.txt.
The third field of /proc//stat is the state. What letter would a process that keeps using the CPU have?
Measure cumulative CPU time
Write the cumulative CPU time (seconds) of the first process to /root/load/cpu_sec.txt as an integer. It must be 5 seconds or more.
Add utime and stime from /proc//stat and divide by the clock tick (usually 100).
Read the PSI pressure metric
In /proc/pressure/cpu, on the some line, take the avg10 value and write it to /root/load/psi.txt. If the file doesn't exist, write unavailable.
The some line of /proc/pressure/cpu has avg10. If the file doesn't exist, write unavailable.
cgroup throttling counter
In the cgroup's cpu.stat, take the nr_throttled value and write it to /root/load/throttle.txt in the form nr_throttled=<값>. If it can't be read, write unavailable.
cpu.stat has nr_periods and nr_throttled. If it can't be read, it is unavailable.
Load judgment script
Create /root/load/loadcheck.sh <임계비율> (the argument is the threshold ratio). If 1분 부하 평균 / CPU 개수 (the 1-minute load average divided by the CPU count) exceeds the threshold ratio, print high load=<값> cpus=<개수> and exit with code 1; if it doesn't, print ok and exit with code 0. With no argument, exit with a non-zero code.
Compare the ratio of the load average divided by the core count with the threshold. You can compare decimals with awk even without bc.