Seven Namespaces and What a Container Really Is
In one line
There is nothing like a struct container in the kernel source. A container is the name people give to a state in which
namespaces + cgroups + a security layer + a rootfs are bundled onto one process, and the one question
namespaces answer is this: what can this process see?
Why this matters
Because of the phrase "create a container", many people imagine that somewhere in the kernel there is a box called a container.
But if you run ps on the host, the processes inside the container show up as they are, and the kernel version is
the same as the host's. If there were a box, this could not happen. What actually happens is the opposite. The process
is simply on the host from the start, and the kernel merely shows it a different version of some resource
lists. It is not isolation; it is a restricted field of view.
This difference shows up immediately in practice. nproc shows 32 inside a container even though it can actually use only 0.5 cores,
two containers each open port 80, a file deleted inside a container remains on the host — all of these can be predicted
only if you know "what is separated and what is not".
How it works
There are seven kinds of namespaces.
| Type | What it separates |
|---|---|
| pid | The process number space. Becoming PID 1 brings the responsibility for signal handling and reaping zombies |
| net | Interfaces, routing tables, iptables rules, sockets, and the port number space — all of them |
| mnt | The mount table. The reason the image rootfs appears as / |
| uts | hostname |
| ipc | System V IPC, POSIX message queues |
| user | UID/GID mappings. The core of rootless containers |
| cgroup | Makes its own cgroup location appear as the root |
It matters that net separates even the port number space. That is why two containers can each open port 80 without conflict. Port conflicts arise not between containers but at the moment you publish to the host.
Each namespace has an inode number, and if you readlink /proc/<pid>/ns/<종류> (where the last part is the namespace type),
that number comes out. The same number means the same namespace.
Measured in practice, it looks like this.
호스트 pid:[4026531836] net:[4026531840] user:[4026531837] time:[4026531834]
alpine pid:[4026532500] net:[4026532562] user:[4026531837] time:[4026531834]
Only user and time are the same. This is because Docker does not use a user namespace by default. So root inside the container has the same UID 0 as root on the host, and this is the root of the security problem covered in the next course.
What it looks like in the field
A Kubernetes Pod uses exactly this structure. In a Pod, a container called pause, about 1 MB in size, starts first, and this container
holds the net/ipc/uts namespaces. The app containers join those namespaces instead of creating new ones,
and only mnt and cgroup are kept separately per container. That is why containers in the same Pod
communicate over localhost and share ports but each has a different filesystem.
You can build this structure yourself with Docker. --network container:<이름> (where the placeholder is the container name) is exactly the command that says
"join someone else's net namespace". It is what you use to reproduce the sidecar pattern with Docker alone.
When does a namespace disappear?
A namespace is not an object but something that exists only while references to it remain. When the last process in it ends and nothing is holding it, the kernel reclaims it. Knowing this property explains a few things in container operations that look strange.
The Pod is gone but the network configuration remains. That means something is still holding that
namespace. There are three ways to hold one: a process running inside it,
a file descriptor with /proc/<PID>/ns/<종류> open (again the last part is the namespace type), and a bind mount placed on
that path. The last is the method ip netns uses, so even after all processes have ended, if the name remains the namespace remains.
You deleted the container but its name keeps lingering. It is the same reason, and in practice it is usually because the cleanup procedure did not release that bind mount.
Conversely, this property is also used on purpose. You create a namespace with no processes in it beforehand and join things to it later. The pause container we saw earlier does exactly that job. Even if an app container dies and comes back up, the network namespace stays alive and the IP does not change. This is why a Pod's IP survives container restarts.
One more point. Namespaces do not divide resources. They decide only what can be seen, and how much can be used is decided by cgroups. So namespaces alone cannot stop one container from using all the CPU. A container looks isolated because both are applied together, and if you apply only one of them, half of it leaks out as is.
What you will do in the next lab
You compare the namespace inodes of the host and a container yourself, check who PID 1 is, and confirm one by one which namespace causes the hostname, the file tree, and the interface list to differ. At the end, you make two containers share a net namespace to reproduce the structure of a Pod by hand.