Diagnosing CPU and Memory Leaks
Break Down a Report of "It's Slow"
Goal
You break "the server is slow" down into which of CPU, memory, or I/O it is. You read the numbers yourself and tell a leak from a cache by the slope.
Why it matters
The most common response to a slowness report is "let's upgrade the specs". But if you add CPU to a machine with a load of 20 because 20 processes are in D state, nothing happens. Finding the numbers that connect symptom to cause is the whole point of this lab.
The constraints of this environment are the training
This Pod has all kernel capabilities removed, so you cannot use perf, strace, or bpftrace.
Instead you read /proc and /sys/fs/cgroup directly. That is what those tools do
internally, and it is also what you end up doing on a customer server where you cannot
install tools.
Steps
- 60-second triage →
/root/perf/01-triage.txt - Read
/proc/stattwice and compute CPU utilization →/root/perf/02-cpu.txt - The top CPU process →
/root/perf/03-top.txt - RSS, PSS, Private_Dirty →
/root/perf/04-mem.txt - Create a leak and confirm it by the slope →
/root/perf/05-slope.txt - cgroup limits →
/root/perf/06-cgroup.txt - Separate anon and file →
/root/perf/07-split.txt - A three-line conclusion →
/root/perf/08-verdict.md
Notes
- It is normal for
top's %CPU to exceed 100 — it is relative to a single core. - The value to look at in
freeisavailable, notfree. - When you add up the memory of several processes, use PSS, not RSS.
60-second triage
60-second triage → /root/perf/01-triage.txt
Write four lines to /root/perf/01-triage.txt: the three loadavg numbers / the count of processes by state (R, D, S) / the number of cores / available memory. The commands are cat /proc/loadavg, ps -eo state --no-headers | sort | uniq -c, nproc, and free -m. This step decides what to look at.
Calculate CPU utilization by hand
Read /proc/stat twice and compute CPU utilization → /root/perf/02-cpu.txt
Read the cpu line of /proc/stat twice, one second apart, compute the utilization, and write only the number (0–100) to /root/perf/02-cpu.txt. busy=user+nice+system+irq+softirq+steal and idle=idle+iowait. This is the calculation top does.
Find who is using the CPU
The top CPU process → /root/perf/03-top.txt
Use ps -eo pid,pcpu,comm --sort=-pcpu | head and write the name of the top process on one line in /root/perf/03-top.txt. If there is no load, start one with python3 -c 'while True: pass' & and look at that.
Tell RSS, PSS, and Private_Dirty apart
RSS, PSS, Private_Dirty → /root/perf/04-mem.txt
Pick any process and save the output of grep -E '^(Rss|Pss|Private_Dirty):' /proc/<PID>/smaps_rollup to /root/perf/04-mem.txt as it is. All three lines must be present.
Create a leak and catch it by the slope
Create a leak and confirm it by the slope → /root/perf/05-slope.txt
Run a program that leaks on purpose, for example python3 -c "import time;L=[];\nwhile True: L.append(bytearray(1024*1024)); time.sleep(0.2)" &. Sample that PID's Private_Dirty four times at 5-second intervals and save one value per line to /root/perf/05-slope.txt. The values must increase monotonically.
See the container's real limit
cgroup limits → /root/perf/06-cgroup.txt
Write the three values memory.max, memory.current, and cpu.max to /root/perf/06-cgroup.txt. In cgroup v2 they are directly under /sys/fs/cgroup/. The point is that these numbers differ from the host's free.
Heap or cache?
Separate anon and file → /root/perf/07-split.txt
Save the output of grep -E '^(anon|file) ' /sys/fs/cgroup/memory.stat to /root/perf/07-split.txt. Even if memory.current reaches the limit, if most of it is file, it is cache and can be reclaimed; if it is anon, the pressure is real.
The conclusion in three lines
A three-line conclusion → /root/perf/08-verdict.md
Three lines in /root/perf/08-verdict.md. (1) Whether the symptom is CPU, memory, or I/O (2) The numbers that were the basis for that judgment (3) What to check next. A conclusion without numbers is a guess.