TT Lab
Get started
Learn Learning paths Courses

Container Security

root Inside the Container Is root on the Host

Continue in TT Lab

In one line

If you have not enabled user namespaces, root in a container is the same UID 0 as root on the host. The claim "containers are isolated, so it is fine to be root inside" is wrong.

Why this was needed

Most official images still start as root. Run docker run --rm node:22-bookworm-slim id and you get uid=0(root). That alone is not surprising, but the next line is the problem.

docker run --rm -v /etc:/host-etc alpine:3.20 sh -c 'echo "# injected" >> /host-etc/hosts'

This works as written. The host's /etc/hosts is modified. If a container were a box, this would be impossible, but as you saw in the earlier course, a container is not a box. It is a process with a restricted view, and on any path opened up by a mount, that process's UID 0 privileges apply as they are. Isolation has not been broken; it was designed to work this way from the start.

How it works

The standard approach is to make the image itself run as an unprivileged user.

RUN useradd --uid 10001 --create-home --shell /usr/sbin/nologin app
COPY --chown=10001:10001 . /app
USER 10001:10001

There is a reason the USER must be a numeric UID rather than a name. When the image's USER is a name, Kubernetes' runAsNonRoot cannot tell whether it is root or not, and it rejects the Pod.

Error: container has runAsNonRoot and image has non-numeric user (app),
cannot verify user is non-root

An unprivileged user cannot open ports below 1024. This is where you may be tempted to add the NET_BIND_SERVICE capability, but it is usually much simpler to change the app port to 8080 and expose it as 80 at the Service layer. Leave the choice of adding privileges for last.

Capabilities split root's privileges into about 40 pieces, and the container default set contains 14 of them (CapEff: 00000000a80425fb). It includes cap_chown, cap_dac_override, cap_fowner, cap_setuid, cap_setgid, cap_net_bind_service, cap_net_raw, cap_sys_chroot, cap_mknod, cap_setfcap, and others, while SYS_ADMIN, NET_ADMIN, and SYS_PTRACE are left out.

The recommended combination looks like this.

--cap-drop=ALL --cap-add=NET_BIND_SERVICE
--security-opt no-new-privileges:true
--read-only --tmpfs /tmp
--user 10001:10001

The default seccomp profile blocks about 44 of the 400-odd system calls. Among them are open_by_handle_at, which was actually used in a container escape, as well as keyctl and kexec_load.

What it looks like in the field

--privileged is a switch that turns off all of the protections above at once. And mounting the Docker socket amounts to the same thing. With that socket, you can start a new privileged container with the host root filesystem mounted. CI runner configurations often leave these two on out of habit, and the moment they do, the rest of the security settings become decoration.

Real incidents follow this same pattern. CVE-2019-5736 overwrote the host's runc binary through /proc/self/exe, and CVE-2024-21626 let an attacker reach the host filesystem with nothing more than a manipulated WORKDIR, because runc did not close a file descriptor. Both were runtime bugs, not kernel bugs, but they worked because the container sits under the same kernel as the host.

How far can you actually cut privileges

"Not running as root" is only the start. What you can cut back in a container comes in four layers: user, privileges (capabilities), filesystem, and system calls. If you touch all four layers, what an attacker can do after a compromise shrinks greatly.

securityContext:
  runAsNonRoot: true
  runAsUser: 10001
  allowPrivilegeEscalation: false     # setuid 로 권한을 되찾는 길을 막는다
  readOnlyRootFilesystem: true
  capabilities:
    drop: ["ALL"]
  seccompProfile:
    type: RuntimeDefault

allowPrivilegeEscalation: false quietly does a lot of work. If it is left on, a setuid binary inside the container can raise privileges again. If you lower the user but do not turn this off, you have done only half the job.

Drop all capabilities and add back only what you need. If you must open a port below 1024, add just NET_BIND_SERVICE. A better answer is to listen on 8080 from the start and have the Service map it to 80, which removes any reason to add privileges at all.

A read-only root reveals where writes are needed. When you turn it on, things usually break at first, and the places where they break are where the program writes. If you attach an emptyDir only at /tmp and the cache path, nothing else can be written. The real value of this setting is that the attacker has nowhere to download and store tools.

The most common stumbling block is file ownership. If the image was built with root ownership and you run it as a different user, it cannot read the files. It is better to match them in advance with COPY --chown=10001:10001 when building the image. For volumes, fsGroup changes the group at mount time.

The default seccomp profile is often enough on its own. RuntimeDefault blocks dozens of dangerous system calls such as unshare and ptrace. A custom profile is more precise, but it carries a maintenance cost: it breaks silently when the kernel or runtime changes. Turning on the default alone gets you most of the benefit.

What you will do in the next lab

You confirm that the default image starts as root, create unprivileged execution in two ways, with a runtime flag and with a Dockerfile, and then apply a read-only root, removal of all capabilities, and blocking of privilege escalation one at a time. At the end, you build an image that meets all of these conditions yourself.