LFCS — Linux Foundation System Administrator
Why You Cannot Create a File With Space to Spare
In one line
Storage is stacked in four layers: block device → partition → filesystem → mount. The first thing to do when something fails is to decide which layer the symptom belongs to, and that judgment alone cuts the candidates to a quarter.
Why this exists
df -h says 35GB is left, yet a single touch fails with No space left on device. It is easy to conclude "the monitoring is off", but df honestly reported only block usage. This error is errno 28, and there are several paths by which the kernel returns this value besides running out of blocks.
The most common is inode exhaustion. A file uses one inode, which holds metadata, separately from its data blocks. On ext4 the number of inodes is fixed at filesystem creation time, so when a huge number of very small files pile up, the inodes run out before the blocks do. You check with df -i, and if IUse% is 100, you have your answer. You cannot increase it, so the only options are deleting files or recreating the filesystem. Session files, backed-up mail queues, and temporary files that are never cleaned up are the usual causes.
The next candidates are worth knowing too. A file that was deleted but is still open — the directory entry is gone so du cannot find it, but a process has it open so the blocks are not returned. It is the most common reason df and du disagree. Reserved blocks — the ext family by default keeps a certain percentage for root only. Mount shadowing — if you mount another filesystem over a directory that has files in it, the files underneath become invisible while still occupying space.
And there is one check worth spending 30 seconds on before diagnosing. Are you really writing to that filesystem? The most common mistake is to run df -h with no arguments, skim the list by eye, and guess. If you directly ask which device a path really belongs to with findmnt -T <경로> (the placeholder is the path), you immediately catch the situation where it is not / but the separate partition /var that is full. Environments where /tmp is a tmpfs are common too, and when that fills up, the same error occurs regardless of the disk, while consuming that much physical memory.
How it works
The six fields of /etc/fstab.
| Number | Field | Description |
|---|---|---|
| 1 | Device | UUID, LABEL, or device path |
| 2 | Mount point | The directory to attach to |
| 3 | Filesystem type | ext4, xfs, nfs, and so on |
| 4 | Options | defaults, noatime, nofail, and so on |
| 5 | dump | For backup tools. Usually 0 |
| 6 | fsck order | 1 for root, 2 for the rest, 0 for no check |
The reason to use a UUID is that device names are not stable. /dev/sdb is a name assigned in the order the kernel discovered devices, so if you add a disk or the controller changes, a different disk can take that name. Then fstab stays syntactically fine but mounts the wrong device. A UUID is engraved in the filesystem, so it follows the device wherever it is plugged in.
Among the options, nofail is especially practical. Without it, boot drops into emergency mode when that device is absent. It is usually added for external disks and network storage entries.
The problem LVM solves. A partition draws a fixed boundary on a disk, so to grow it later you need free space adjacent behind it. LVM gathers physical volumes (PVs) into a pool called a volume group (VG), and cuts logical volumes (LVs) out of that pool. So you can create one volume across several disks, and after adding a disk to grow the pool, you can extend the volume online. Snapshots come from this layer too.
The RAID level is chosen by requirements. RAID 0 gives only performance and capacity with no redundancy at all (if one disk dies you lose everything). RAID 1 is a straight mirror so it is safe but capacity is halved. RAID 5 survives one failure with a single parity and has good capacity efficiency, but it has to read all the remaining disks during a rebuild, so if a second failure happens then, it is over. The larger the disks, the longer the rebuild, and the bigger this risk gets. RAID 10 stripes mirrors, gaining write performance and rebuild safety at the cost of half the capacity. And no RAID is a backup. RAID is a device that survives hardware failure; it does not bring back a file deleted by mistake.
What it looks like in the field
In the case the author experienced, Avail was 35GB yet a single 0-byte file could not be created. If you check in order, the cause comes out in 5 minutes — pin down the filesystem the path belongs to, look at inodes with df -i, look for deleted open files, and check reserved blocks and mount shadowing. Memorizing this order is very valuable. Because the same error message comes from several causes, guessing from the message alone leaks you down a different path each time.
The first candidate when df and du disagree is also always the same. A file that was deleted but is held by a process. The safest way to release it is not restarting the service but sending a signal that makes it reopen its log file, and the next resort is to empty the descriptor directly. The reason to keep this order is that a restart wipes out all the diagnostic information of that moment.
What you will check in the next quiz
In this lab environment, commands such as mount, umount, mkfs, fdisk, parted, pvcreate, and swapon are all blocked. It is a container with kernel capabilities removed, so you cannot actually create or attach a filesystem. So this module concentrates on concepts and judgment criteria. To be honest, storage skill on the LFCS grows only by doing it by hand — so the very next module (a lab where you do LVM, mkfs, mount, swap, NFS, and autofs yourself on a real VM) is the place for that practice. Here you firmly establish the layered structure, the difference between df and du, the diagnostic order for inode exhaustion, the meaning of each fstab field and the reason for UUIDs, and the criteria for choosing between LVM and RAID, and check them with a quiz.