The Five Faces of "No space left on device"
In one line
ENOSPC does not mean "there are no blocks" but "the kernel has no resource to accommodate this write." That resource may be blocks, inodes, or free space excluding the reserved portion.
Why this was needed
df -h /var shows 35G free, yet touch /var/log/test fails with No space left on device. The application can't write logs, the deployment dies while creating temporary files, and the DB can't write its WAL.
Concluding here that "df is wrong" gets you nowhere. df honestly reported what it knows (block usage); the cause of the failure simply wasn't blocks.
How it works
Ordered by probability, there are five possibilities.
0. Mistaking the path. / has free space but /var is a separate partition and that one is full. It is confirmed in 30 seconds - first establish which filesystem the path belongs to with findmnt -T /var/log/app.log. If /tmp is tmpfs, you get ENOSPC regardless of the disk, and at the same time it eats that much memory.
1. Inode exhaustion. ext2/3/4 fix the number of inodes at format time and cannot increase it afterward. The default profile is roughly one per 16KB, so a 100GB filesystem has about 6.55 million. For a workload whose average file size is smaller than that (session files, mail queues, caches), inodes run out before blocks. If df -i shows IUse% at 100%, it is confirmed. XFS allocates dynamically, so it virtually has no such problem.
2. Deleted but still open files. This is the structure you learned in the earlier course. rm removes only the link, and blocks are returned only when the link count is 0 and the open fd count is 0. du walks paths, so it cannot count that file, which has no directory entry. The gap between df and du is the decisive signal. A regular cause is logrotate not sending the reopen signal after rotation.
3. Reserved blocks. ext-family filesystems by default leave 5% for root only. There are two purposes - to leave room for the administrator to clean up even when it is full, and to give the allocator room to find contiguous blocks and so delay fragmentation. The symptom is distinctive. It works as root but fails only for the service account.
4. Mount shadowing. If you mount another filesystem over a directory that has files in it, the original files are no longer visible but still occupy space on the filesystem underneath. du walks only the upper one, so no tool shows it.
5. inotify. When the watch limit is exceeded, inotify_add_watch returns ENOSPC. It has nothing to do with the disk, but the wording is the same.
What you see in the field
df 20G / du 3.1G. When there is a 17GB difference, suspect deleted open files. In lsof -nP +L1, they are the entries whose NLINK column is 0, and SIZE is the bytes not yet returned. On a minimal image without lsof, you sweep /proc/*/fd yourself to find (deleted). That is exactly the method you use in the lab.
The pitfall of emergency treatment. truncate -s 0 /proc/PID/fd/N empties the file and returns the space immediately. But if the file is not opened with O_APPEND, the process keeps writing at its old offset and the file becomes sparse, and the size in ls -l still looks large. It is for emergency use only; the proper fix is a reopen signal (such as kill -USR1) or a restart.
When diagnosis makes the outage worse. Running a full find / scan on a disk whose I/O is already saturated makes things worse. Fix the filesystem boundary with -xdev and start with a narrowed scope.
When df and du say different things
You get an alert that the disk is full and log in, and no matter how much you add up with du, it doesn't
come to that amount. It is one of three things.
A file that was deleted but is still open. Deleting a file in Unix only detaches the name from the directory.
As long as some process has the file open, the blocks are not returned. The typical case is when you deleted a log file with rm
but the process keeps writing to that fd. du walks along names, so it can't see this usage,
while df counts blocks, so it can.
lsof +L1 # 링크 수가 0인, 즉 지워졌는데 열려 있는 파일
ls -l /proc/<pid>/fd | grep deleted
The way to get it back is to restart the process, and if you can't, you truncate the fd.
: > /proc/<pid>/fd/3 doesn't delete the file but only sets its size to 0.
This is why you use, instead of rm, truncate or logrotate's copytruncate.
When inodes run out first. Blocks remain but you can't create files.
The same message, No space left on device, appears, so the two can't be told apart from
the symptom alone. The causes are cache directories, mail queues and session files that pile up millions of small files.
df -i # 사용률이 100%인 파일 시스템
find /var -xdev -type f | cut -d/ -f1-4 | sort | uniq -c | sort -rn | head
If you leave out -xdev, it crosses into other mounts and counts the wrong place. The number of inodes is fixed
when the filesystem is created (for ext4), so it can't be increased later. XFS allocates them dynamically,
but has an imaxpct upper limit.
A file hidden by a mount. If you write a lot of files into /data and later mount a disk on that spot,
the files that were there become invisible while still occupying space on the root filesystem.
You can see them only after you take the mount away.
mkdir /mnt/root && mount --bind / /mnt/root
du -shx /mnt/root/data
Finally, remember the reserved blocks. ext4 by default keeps 5% for root. To ordinary users
it already looks full, but df says there is still room. On a data-only disk you may reduce it with tune2fs -m 1, but you
don't touch it on the root filesystem — that margin is the last place the system has to recover itself.
What you will do in the next lab
You read the inode usage and directory sizes yourself and find out why 500 empty files consume space. Then you create a deleted-but-open file yourself, find the fd that holds it in /proc, and recover the contents through that fd path. At the end you write a script that finds such ghost files across the whole system.