TT Lab
Get started
Learn Learning paths Courses

Operating Systems

Filesystems — A Design That Separates the Name From the Thing

Continue in TT Lab

In a nutshell

The Unix filesystem separated a file's substance (the inode) from its name (the directory entry), and that one decision explains behaviors such as hard links, deleting an open file, and atomic replacement.

Why this was needed

A file has two kinds of information: its contents, and its metadata (size, permissions, owner, modification time, and which blocks hold the data). It also has a name. Bundling all three into one lump is simple, but calling the same file by two names or renaming it becomes expensive.

The Unix choice was to keep the metadata and contents in the inode and the name in the directory. A directory is a special file, and its contents are a list of "name → inode number" pairs.

How it works

An inode holds the file size, permissions, timestamps, the link count, and pointers to the data blocks. The pointer structure is interesting. The first few point directly at data blocks, the next one is a single indirect pointer to a block of pointers, followed by double and triple indirect ones. A small file needs only direct pointers and is fast, and large files can still be represented. The price is that the deeper the levels, the more block reads an access needs.

Some consequences follow naturally from this structure.

A hard link is creating one more name that points to the same inode. There is no distinction between the original and the copy; the two are completely equal. The inode's link count just goes up by one.

Deleting a file is actually deleting a name (which is why the system call is named unlink). The blocks are reclaimed only when the link count reaches 0 and no process has the file open. That is how a situation arises in which you delete a log file but the free disk space does not increase. It is because a process still has the file open, and the space comes back only when you restart that process or close the file descriptor.

Renaming (rename) within the same filesystem only modifies a directory entry, so it is atomic. The standard way to safely replace a configuration file comes from this. You write the new contents to a temporary file, flush it to disk with fsync, and then swap it in with rename. A reader always sees either the old contents or the new contents, and never sees a half-written file.

Journaling is a safeguard against crashes. If you first record a metadata change in the journal before writing it to its actual location, then even if the power goes out midway, the journal can be replayed at reboot to restore consistency. The default mode of ext4 is ordered, which journals only metadata. Journaling data as well is safe, but every write happens twice.

What it looks like in the field

There are cases where df says there is free space but writes fail. It is one of two things. Either the inodes have run out (check with df -i), or someone still has a deleted file open and the space has not been reclaimed. For the latter, you can find open files whose link count is 0 with lsof +L1. The typical situation is one in which log rotation is misconfigured, the old file is deleted, and the application keeps writing to that descriptor.

Did the write really reach the disk

The procedure above for safely replacing a configuration file included fsync. To understand why this one step is necessary, you need to know the layers a write passes through.

When a program calls write, the data goes into the kernel's page cache, and at that moment the call returns success. Nothing has gone to disk yet. The kernel sends it down later on its own, and if the power goes out in between, that data is as if it never existed. fsync is a request to "send it down now and wait until it is finished".

There is one more thing people often miss here. Running fsync on a file's contents does not make the file's name safe. If you create a new file, write it, run fsync, and then the power goes out, the contents may remain but the directory entry may be missing, leaving no way to reach that file. That is why, after creating a new file or after a rename, you must fsync the directory as well. This is exactly what databases and logging systems do.

And the layers do not end at the kernel. The disk itself also has a write cache, so even when the device answers "written", the data may still be in a volatile buffer. A proper storage device supports a cache flush command and the filesystem sends it, but some cheap devices accept that command, do nothing, and answer that it succeeded. On such devices, no software can guarantee crash safety.

The practical conclusion is simple. Use fsync for data you cannot afford to lose, and accept being slower as the price. Conversely, do not use it for caches or temporary files that can be recreated at any time. If you do not distinguish the two kinds and fsync everything, performance collapses, and if you omit it everywhere, then when the power goes out one day you will not even know what disappeared.

What you will do in the lab that follows

You will build everything described here by hand. You will make scripts that tell whether two files are the same by inode, read a file from which the name has been removed, and find the process that is holding space on a file that was deleted. You will also make a script that swaps in a configuration file. The grader will have that file open while it runs your replacement, to check whether a reader sees a half-written file.