TT Lab
Get started
Learn Learning paths Courses

Diagnosing CPU and Memory Leaks

What RSS, VSZ and PSS Each Measure

Continue in TT Lab

In one line

VSZ is a promise, RSS is the physical memory actually resident, and PSS is your share after dividing up the shared pages. Adding up the RSS of several Pods gives more than what is really used — that is when you should look at PSS.

Why this matters

Half of the reports that say "it looks like a memory leak" are not leaks. A process can grow for several reasons.

The four call for different responses. If you skip telling them apart and just say "let's add memory", the same thing happens again a few days later.

How it works

The most accurate way to see one process's real usage is a single smaps_rollup line.

grep -E '^(Rss|Pss|Private_Dirty|Swap):' /proc/<PID>/smaps_rollup
Field Meaning
Rss Everything resident in physical memory (including shared libraries)
Pss The shared portion divided by the number of users — the only value to use when you add things up
Private_Dirty The part only this process uses and that cannot be written back to a file — use this when looking for leaks

Tell a leak from a cache by the slope, not a value at one moment. With a steady load applied, sample Private_Dirty every few minutes: if it increases monotonically, it is a leak; if it flattens out at some level, it is a cache.

Common misconceptions

Being alarmed by VSZ. A 64-bit JVM or Go runtime reserves tens of GB of VSZ. It has only reserved address space, not physical memory. You can almost always ignore VSZ.

Adding up RSS. If 10 containers use the same libc, those pages are physically one copy, but RSS counts them 10 times. Use PSS for totals.

Worrying that the free value in free -h is small. Linux uses all spare memory as page cache. The value to look at is available.

Where to see a container's memory limit

The host's free or top does not know the share allocated to a container. Read the cgroup directly.

cat /sys/fs/cgroup/memory.max        # 한도 (max 면 제한 없음)
cat /sys/fs/cgroup/memory.current    # 지금 쓰는 양
cat /sys/fs/cgroup/memory.stat       # 항목별 내역

memory.current includes the page cache. A process that reads many files looks close to the limit because of the cache, but when pressure comes the kernel drops the cache first, so no OOM occurs. That is why "we are using 90% of memory" is not necessarily a danger.

Whether it is truly dangerous is shown by two lines of memory.stat.

anon      1234567890     ← 익명 메모리. 이건 버릴 수 없다
file       987654321     ← 페이지 캐시. 압박이 오면 버려진다

If anon is close to the limit, an OOM is imminent.

How to check whether an OOM happened

If a Pod suddenly restarted and the logs show nothing, it is usually an OOM. The application has no way to know it is about to die, so it cannot leave any last words.

kubectl describe pod <파드> | grep -A3 "Last State"
    Last State:     Terminated
      Reason:       OOMKilled
      Exit Code:    137

# cgroup 쪽 기록
cat /sys/fs/cgroup/memory.events
oom 3
oom_kill 1

oom is the number of times the cgroup hit the limit and tried to reclaim memory, and oom_kill is the number of times it actually killed something. If only oom increases and oom_kill stays 0, it means nothing has died yet, but the cgroup is under constant pressure, and from then on latency gets worse. If you miss this signal, it only looks like "sometimes slow".

What to tune for each language

Lowering the container limit does nothing if the runtime does not know about it.

Runtime What to tune If you do not
JVM -XX:MaxRAMPercentage=75 Older JVMs size the heap based on host memory
Node.js --max-old-space-size=<MB> The default heap can be larger than the container limit
Go GOMEMLIMIT GC runs late and the process exceeds the limit
Python No separate setting Control it with the number of workers (gunicorn -w)

GOMEMLIMIT has existed since Go 1.19, and if you set it to about 90% of the limit, GC runs more often around that line. If you put a Go service in a tight limit without it, GC takes its time and the service runs into an OOM.

What really matters in practice

A container dies from the cgroup limit, not from the host.

cat /sys/fs/cgroup/memory.max        # 한도
cat /sys/fs/cgroup/memory.current    # 지금
cat /sys/fs/cgroup/memory.events     # oom_kill 횟수

memory.current also includes page cache. That is why a process that reads many files hits the limit even when its real heap is small. You have to look at anon (anonymous memory) and file (cache) in memory.stat separately to see the real cause. An OOM kill is done by the kernel, so nothing is left in the application log — often the only evidence is whether oom_kill in memory.events has increased.