The Four Doors That Open a Container
In one line
Container escapes usually come not from kernel vulnerabilities but from settings we opened ourselves. If you know four patterns, you can catch them in review in 30 seconds.
This lesson is about defense. For each item, it pairs "why it is dangerous" with "what blocks it".
Why this was needed
Checking that a container is non-root is not enough to prove the host boundary. If you open the Docker socket, the host filesystem, the PID namespace, or excessive capabilities, even a low-privilege process inside the container can obtain host privileges by bypassing the restrictions. So you must review not only the image user but also the resources and kernel privileges attached at run time.
Door 1 — Mounting the Docker socket
volumes:
- /var/run/docker.sock:/var/run/docker.sock
This is a line commonly seen in CI runners and monitoring agents. This socket is the Docker daemon's API, and the daemon runs as root on the host. If you can access the socket, you can start a new privileged container with the whole host root mounted. That is, this one line is effectively the same as handing out host root privileges.
How to block it — Do not give the socket to the container. If you need to build containers, use a builder that does not need a daemon (kaniko, buildah), and if the goal is to look up information, put a read-only proxy in front and restrict the permitted endpoints with an allowlist. This is why LabHub's image builds use kaniko.
Door 2 — --privileged
--privileged gives all capabilities, opens device access, and effectively
disables seccomp/AppArmor. In this state, it is possible to open the host disk as
/dev/sda and mount it.
How to block it — Give only the capabilities you need, individually. For example, if the goal is low-port
binding, NET_BIND_SERVICE alone is enough. On Kubernetes, reject privileged Pods outright
with Pod Security Admission at baseline or above.
Door 3 — Mounting a host path
volumes:
- /:/host # 최악
- /etc:/etc-host # 나쁨
- /var/log:/logs # 상황에 따라
Not to mention /, even /etc alone reaches shadow, sudoers, and the cron files.
If it is writable, it is a host compromise as it stands.
How to block it — Narrow the mount path to the minimum scope and add :ro.
Use a named volume instead of a host path, and on Kubernetes it is more reliable
to forbid hostPath itself with a policy engine.
Door 4 — Shared namespaces
--pid=host makes every process on the host visible, and --net=host uses the
host's network stack as it is. The former opens a way to reach other processes' filesystems through /proc/<pid>/root, and the latter lets you reach management ports that were
opened only on localhost.
How to block it — Keep the default (isolation). Even in cases where it is truly needed, such as monitoring, narrow it with read-only plus minimal capabilities.
At a glance
| Setting | Practical meaning | Alternative |
|---|---|---|
docker.sock mount |
Host root | kaniko/buildah, read-only proxy |
--privileged |
All capabilities + devices | Only the needed caps, PSA baseline |
-v /:/host |
Host filesystem | Narrow scope + :ro, named volume |
--pid=host |
Access to other processes | Keep default isolation |
And the layers of defense
Even with good settings, runtime vulnerabilities remain. Stack up layers.
- UID —
USER 10001. If you do not run as root, most paths are blocked. - capability — Drop all and add only what you need.
- NO_NEW_PRIVS — Cut off the route of raising privileges through setuid.
- Read-only root —
--read-onlyplus tmpfs only where needed. - seccomp — Block dangerous system calls at the kernel's entrance.
LabHub's lab Pods apply all five. So whatever you do as root inside a lab, it does not reach the host.
What it looks like in the field
- A CI runner is mounting the socket → anyone who can commit to the repository can take over the host.
- "Let's try it with privileged first and cut it down if it works" → the commit that cuts it down never comes.
- A log collector attached
/var/logas rw → it can delete the audit logs.
What to look for in the next check
The quiz that follows separates the paths opened by the Docker socket, host mounts, and PID sharing from the roles of minimal capabilities, NoNewPrivs, and seccomp, which block them. Using the kernel privilege values you confirmed in the earlier lab as evidence, filter out the wrong answer that a single defense is sufficient.