TT Lab
Get started
Learn Learning paths Courses

Diagnosing CPU and Memory Leaks

Break Down a Report of "It's Slow"

Continue in TT Lab

Goal

You break "the server is slow" down into which of CPU, memory, or I/O it is. You read the numbers yourself and tell a leak from a cache by the slope.

Why it matters

The most common response to a slowness report is "let's upgrade the specs". But if you add CPU to a machine with a load of 20 because 20 processes are in D state, nothing happens. Finding the numbers that connect symptom to cause is the whole point of this lab.

The constraints of this environment are the training

This Pod has all kernel capabilities removed, so you cannot use perf, strace, or bpftrace. Instead you read /proc and /sys/fs/cgroup directly. That is what those tools do internally, and it is also what you end up doing on a customer server where you cannot install tools.

Steps

  1. 60-second triage → /root/perf/01-triage.txt
  2. Read /proc/stat twice and compute CPU utilization → /root/perf/02-cpu.txt
  3. The top CPU process → /root/perf/03-top.txt
  4. RSS, PSS, Private_Dirty → /root/perf/04-mem.txt
  5. Create a leak and confirm it by the slope → /root/perf/05-slope.txt
  6. cgroup limits → /root/perf/06-cgroup.txt
  7. Separate anon and file → /root/perf/07-split.txt
  8. A three-line conclusion → /root/perf/08-verdict.md

Notes

60-second triage

60-second triage → /root/perf/01-triage.txt

Write four lines to /root/perf/01-triage.txt: the three loadavg numbers / the count of processes by state (R, D, S) / the number of cores / available memory. The commands are cat /proc/loadavg, ps -eo state --no-headers | sort | uniq -c, nproc, and free -m. This step decides what to look at.

Calculate CPU utilization by hand

Read /proc/stat twice and compute CPU utilization → /root/perf/02-cpu.txt

Read the cpu line of /proc/stat twice, one second apart, compute the utilization, and write only the number (0–100) to /root/perf/02-cpu.txt. busy=user+nice+system+irq+softirq+steal and idle=idle+iowait. This is the calculation top does.

Find who is using the CPU

The top CPU process → /root/perf/03-top.txt

Use ps -eo pid,pcpu,comm --sort=-pcpu | head and write the name of the top process on one line in /root/perf/03-top.txt. If there is no load, start one with python3 -c 'while True: pass' & and look at that.

Tell RSS, PSS, and Private_Dirty apart

RSS, PSS, Private_Dirty → /root/perf/04-mem.txt

Pick any process and save the output of grep -E '^(Rss|Pss|Private_Dirty):' /proc/<PID>/smaps_rollup to /root/perf/04-mem.txt as it is. All three lines must be present.

Create a leak and catch it by the slope

Create a leak and confirm it by the slope → /root/perf/05-slope.txt

Run a program that leaks on purpose, for example python3 -c "import time;L=[];\nwhile True: L.append(bytearray(1024*1024)); time.sleep(0.2)" &. Sample that PID's Private_Dirty four times at 5-second intervals and save one value per line to /root/perf/05-slope.txt. The values must increase monotonically.

See the container's real limit

cgroup limits → /root/perf/06-cgroup.txt

Write the three values memory.max, memory.current, and cpu.max to /root/perf/06-cgroup.txt. In cgroup v2 they are directly under /sys/fs/cgroup/. The point is that these numbers differ from the host's free.

Heap or cache?

Separate anon and file → /root/perf/07-split.txt

Save the output of grep -E '^(anon|file) ' /sys/fs/cgroup/memory.stat to /root/perf/07-split.txt. Even if memory.current reaches the limit, if most of it is file, it is cache and can be reclaimed; if it is anon, the pressure is real.

The conclusion in three lines

A three-line conclusion → /root/perf/08-verdict.md

Three lines in /root/perf/08-verdict.md. (1) Whether the symptom is CPU, memory, or I/O (2) The numbers that were the basis for that judgment (3) What to check next. A conclusion without numbers is a guess.