Diagnosing CPU and Memory Leaks
Disks and File Descriptors
In one line
If df shows free space but you cannot write, the inodes have run out, or someone still has a deleted file open. Neither is visible with df -h.
Why this matters
These are the two most common causes behind disk-related reports.
Inode exhaustion. Every file uses one inode. With millions of small files, the inodes run out first even when capacity remains. You only see it with df -i.
Deleted but still open files. If you rm a log file while a process still has it open, Linux removes only the link and does not return the actual blocks. df still says the disk is full, and du says the file is not there. This mismatch is the decisive clue.
# du 와 df 가 다르면 이걸 본다
ls -l /proc/*/fd 2>/dev/null | grep deleted
The fix is to make the process reopen the file (for logs, copytruncate in logrotate, or SIGHUP to the daemon). Restarting the process also works.
How it works
A file descriptor leak kills slowly.
ls /proc/<PID>/fd | wc -l # 지금 몇 개 열었나
cat /proc/<PID>/limits | grep 'open files' # 한도
If this number increases monotonically over time, some code is not closing files. The moment it reaches the limit, every request fails with Too many open files. Raising the limit buys time; it does not fix anything.
Common misconceptions
The belief that ulimit -n being unlimited is safe. A container's default RLIMIT_NOFILE is effectively unlimited (around a billion), but some old daemons try to allocate a connection table of that size at startup, use up all the memory, and die. Unlimited is not always good.
Two numbers that tell you whether the disk is slow
Of the many columns in iostat -x 1, only two are actually used.
Device r/s w/s rkB/s wkB/s r_await w_await aqu-sz %util
nvme0n1 120 380 4800 15200 0.31 0.52 1.24 38.2
await(r_await, w_await) — The response time (ms) for one request, including the time it waited in the queue. Normal is under 1 ms for NVMe, a few ms for a SATA SSD, and around 10 ms for an HDD. If this jumps to several times its usual value, the disk is the bottleneck.aqu-sz— The average queue length. Above 1, requests are lining up.
Do not trust %util. This value is "the fraction of time at least one request was in progress", so
on an NVMe that handles requests in parallel, there can still be headroom at 100%. It is a metric
from the days of spinning disks.
Which process is using the disk
# 실시간으로 I/O 상위 프로세스
iotop -oPa # -o: 실제로 I/O 하는 것만, -a: 누적
# iotop 이 없을 때 (컨테이너에서 흔하다)
cat /proc/<PID>/io # read_bytes, write_bytes 를 두 번 읽어 차이를 낸다
read_bytes in /proc/<PID>/io is the amount actually read from the disk, and rchar is
the amount handed over through read system calls. A large difference between the two means the page cache is working well,
which is a good sign. Conversely, if the two are similar, the cache is not helping and every read hits the disk.
Limiting I/O in containers
Unlike CPU and memory, Kubernetes has no I/O limit field. cgroup v2 has
io.max, but it is not exposed through the Pod spec. So the practical responses are different.
- Physically separate the noisy neighbor. Send workloads that burst I/O, such as log collectors or backups, to a separate node or a separate disk.
- Throttle it in the application. Apply
ionice -c3(the idle class) to a backup script, or use the limit a tool provides, such asrsync --bwlimit. - Buy a network storage IOPS tier instead of using local disk. In the cloud, this is the most reliable option.
ionice only takes effect with the CFQ and BFQ schedulers. With none (multi-queue), the default on modern NVMe,
it has no effect, so check first with cat /sys/block/nvme0n1/queue/scheduler.
What really matters in practice
The difference between create and copytruncate in log rotation settings comes into play here.
create: Renames the original and creates a new file. The process keeps writing to the old file (now renamed), so you must send SIGHUP to the daemon to make it reopen.copytruncate: Copies the contents and then truncates the original to 0 bytes. The process can keep writing as it is, but logs written between the copy and the truncate are lost.
Neither is free. If the application handles SIGHUP, use create; if it cannot, use copytruncate and be aware of the possible loss.