Diagnosing CPU and Memory Leaks
What RSS, VSZ and PSS Each Measure
In one line
VSZ is a promise, RSS is the physical memory actually resident, and PSS is your share after dividing up the shared pages. Adding up the RSS of several Pods gives more than what is really used — that is when you should look at PSS.
Why this matters
Half of the reports that say "it looks like a memory leak" are not leaks. A process can grow for several reasons.
- A real leak: allocations that are never freed keep piling up
- Cache growth: the application deliberately fills a cache (it stops when it reaches its limit)
- Heap fragmentation: memory was freed but could not be returned to the OS
- Allocator behavior: glibc malloc does not readily return small blocks to the OS
The four call for different responses. If you skip telling them apart and just say "let's add memory", the same thing happens again a few days later.
How it works
The most accurate way to see one process's real usage is a single smaps_rollup line.
grep -E '^(Rss|Pss|Private_Dirty|Swap):' /proc/<PID>/smaps_rollup
| Field | Meaning |
|---|---|
Rss |
Everything resident in physical memory (including shared libraries) |
Pss |
The shared portion divided by the number of users — the only value to use when you add things up |
Private_Dirty |
The part only this process uses and that cannot be written back to a file — use this when looking for leaks |
Tell a leak from a cache by the slope, not a value at one moment. With a steady load applied, sample Private_Dirty every few minutes: if it increases monotonically, it is a leak; if it flattens out at some level, it is a cache.
Common misconceptions
Being alarmed by VSZ. A 64-bit JVM or Go runtime reserves tens of GB of VSZ. It has only reserved address space, not physical memory. You can almost always ignore VSZ.
Adding up RSS. If 10 containers use the same libc, those pages are physically one copy, but RSS counts them 10 times. Use PSS for totals.
Worrying that the free value in free -h is small. Linux uses all spare memory as page cache. The value to look at is available.
Where to see a container's memory limit
The host's free or top does not know the share allocated to a container. Read the cgroup directly.
cat /sys/fs/cgroup/memory.max # 한도 (max 면 제한 없음)
cat /sys/fs/cgroup/memory.current # 지금 쓰는 양
cat /sys/fs/cgroup/memory.stat # 항목별 내역
memory.current includes the page cache. A process that reads many files
looks close to the limit because of the cache, but when pressure comes the kernel drops the cache first,
so no OOM occurs. That is why "we are using 90% of memory" is not necessarily a danger.
Whether it is truly dangerous is shown by two lines of memory.stat.
anon 1234567890 ← 익명 메모리. 이건 버릴 수 없다
file 987654321 ← 페이지 캐시. 압박이 오면 버려진다
If anon is close to the limit, an OOM is imminent.
How to check whether an OOM happened
If a Pod suddenly restarted and the logs show nothing, it is usually an OOM. The application has no way to know it is about to die, so it cannot leave any last words.
kubectl describe pod <파드> | grep -A3 "Last State"
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
# cgroup 쪽 기록
cat /sys/fs/cgroup/memory.events
oom 3
oom_kill 1
oom is the number of times the cgroup hit the limit and tried to reclaim memory, and oom_kill is the number of times it actually killed something.
If only oom increases and oom_kill stays 0, it means nothing has died yet, but the cgroup is under constant pressure,
and from then on latency gets worse. If you miss this signal, it only looks like "sometimes slow".
What to tune for each language
Lowering the container limit does nothing if the runtime does not know about it.
| Runtime | What to tune | If you do not |
|---|---|---|
| JVM | -XX:MaxRAMPercentage=75 |
Older JVMs size the heap based on host memory |
| Node.js | --max-old-space-size=<MB> |
The default heap can be larger than the container limit |
| Go | GOMEMLIMIT |
GC runs late and the process exceeds the limit |
| Python | No separate setting | Control it with the number of workers (gunicorn -w) |
GOMEMLIMIT has existed since Go 1.19, and if you set it to about 90% of the limit, GC runs more
often around that line. If you put a Go service in a tight limit without it, GC takes its time
and the service runs into an OOM.
What really matters in practice
A container dies from the cgroup limit, not from the host.
cat /sys/fs/cgroup/memory.max # 한도
cat /sys/fs/cgroup/memory.current # 지금
cat /sys/fs/cgroup/memory.events # oom_kill 횟수
memory.current also includes page cache. That is why a process that reads many files hits the limit even when its real heap is small. You have to look at anon (anonymous memory) and file (cache) in memory.stat separately to see the real cause. An OOM kill is done by the kernel, so nothing is left in the application log — often the only evidence is whether oom_kill in memory.events has increased.