Mount Namespaces and Propagation Modes
In one line
/proc/self/mountinfo holds "the whole mount tree this process sees," and it also records the propagation type, which the mount command does not show.
Why you need this
Once you start using containers, mounts suddenly get hard. A volume attached on the host is not visible in the container, something mounted inside the container leaks to the host, or you deleted the container but the mount remains. All of these phenomena are explained by mount namespaces and propagation.
How it works
Reading mountinfo
36 25 0:31 / /sys/fs/cgroup rw,nosuid,nodev,noexec shared:9 - cgroup2 cgroup2 rw
[1][2][3] [4] [5] [6] [7] [8] [9] [10] [11]
| Field | Meaning |
|---|---|
| 1 | Mount ID |
| 2 | Parent mount ID |
| 3 | major:minor device number |
| 4 | The root within the source — for a bind mount, this is not / |
| 5 | Mount point |
| 6 | Mount options |
| 7 | Optional fields — propagation (shared:, master:, propagate_from:, private if absent) |
| 8 | Separator - |
| 9–11 | Filesystem type, source, superblock options |
Field 4 is the key to identifying a bind mount. If you attached a subdirectory of the same device, that subpath is written there instead of /.
The four propagation types
| Type | mountinfo display | Behavior |
|---|---|---|
| private | (no marker) | Changes to this mount are propagated nowhere |
| shared | shared:N |
Bidirectional propagation. A mount made here is visible to its peers, and vice versa |
| slave | master:N |
One-way. It receives the master's changes but does not send its own |
| unbindable | unbindable |
It cannot be the source of a bind mount |
This is why container runtimes use rslave or rprivate when they attach volumes. It is a problem if a mount made by mistake inside a container leaks to the host, and conversely there are often cases where what the host mounts must be visible.
tmpfs
A filesystem in memory. It disappears on reboot.
tmpfs /var/cache/app tmpfs size=256M,mode=1777,noexec,nosuid,nodev 0 0
- Always specify
size=. Without it, the default is half of physical memory, and if someone writes a large file, the whole system comes under memory pressure. - tmpfs can be swapped out. If you want a true RAM disk, that is
ramfs, but ramfs has no size limit and is more dangerous. mode=1777is the sticky bit. As with/tmp, everyone can write but cannot delete other people's files.
systemd's .mount units
You can also define a mount with a unit file instead of fstab. The unit name must be the mount point path, escaped.
# /etc/systemd/system/var-cache-app.mount
[Unit]
Description=Application cache (tmpfs)
[Mount]
What=tmpfs
Where=/var/cache/app
Type=tmpfs
Options=size=256M,mode=1777,noexec,nosuid,nodev
[Install]
WantedBy=local-fs.target
/var/cache/app → var-cache-app.mount. Drop the leading slash and turn the remaining slashes into hyphens. If the name does not match the path, systemd rejects the unit. You can get the exact name with systemd-escape --path /var/cache/app.
When a mount disappears or is not visible
Bind mounts and tmpfs look simple, but because of propagation and namespaces they sometimes behave differently from what you expect.
A mount made later inside a container is not visible. Even if you mount a new disk at /mnt/data
on the host, it does not appear in a container that had already bind-mounted that path.
This is because, if the propagation type is private, later mounts are not passed along. That is
why Kubernetes uses mountPropagation: HostToContainer.
findmnt -o TARGET,SOURCE,PROPAGATION /mnt/data
tmpfs uses memory. It looks like a filesystem in df, but it is actually RAM, and
in a container, that memory counts toward the cgroup limit. If you put /tmp on tmpfs and
write a large file, it dies of OOM. Always set the size.
tmpfs /tmp tmpfs size=512M,mode=1777,nosuid,nodev 0 0
If you do not mount onto an empty directory, the original contents are hidden. Those files are not gone; they stay underneath and keep taking up space. To get them back, unmount, or bind the root somewhere else and look there.
mkdir /mnt/under && mount --bind / /mnt/under
ls /mnt/under/mnt/data # 가려진 원래 내용
When umount fails, look at who is using it.
lsof +f -- /mnt/data
fuser -vm /mnt/data
umount -l /mnt/data # 마지막 수단: 이름만 떼고 나중에 정리한다
-l (lazy) stays actually attached until the open files are closed, so it is not
safe in a situation where you are physically removing the disk.
An option applied to a bind mount has to be applied twice. mount --bind -o ro does not
make it read-only on Linux. You have to apply it again after attaching.
mount --bind /src /dst
mount -o remount,ro,bind /dst
What it looks like in the field
You delete the container but the mount remains. If a container creates a mount while the propagation type is shared, it leaks out into the host namespace. Even when the container dies, that mount stays on the host and holds the disk. You have to find it with findmnt and clean it up manually.
Not setting the tmpfs size leads to OOM. Logs piled up in a tmpfs created without size= and ate all the memory. It is easy to forget that it is memory, not disk.
What you will do in the next lab
You parse /proc/self/mountinfo yourself to make a table of the mount tree and propagation types, write the fstab lines for a bind mount and for tmpfs, and create a .mount unit file with the exact name.