root Inside the Container Is root on the Host
In one line
If you have not enabled user namespaces, root in a container is the same UID 0 as root on the host. The claim "containers are isolated, so it is fine to be root inside" is wrong.
Why this was needed
Most official images still start as root. Run docker run --rm node:22-bookworm-slim id
and you get uid=0(root). That alone is not surprising, but the next line
is the problem.
docker run --rm -v /etc:/host-etc alpine:3.20 sh -c 'echo "# injected" >> /host-etc/hosts'
This works as written. The host's /etc/hosts is modified. If a container were a box, this would
be impossible, but as you saw in the earlier course, a container is not a box. It is a process with a restricted view, and on any path opened up by a mount, that process's UID 0 privileges apply as they are. Isolation has not been broken; it was designed to work this way from the start.
How it works
The standard approach is to make the image itself run as an unprivileged user.
RUN useradd --uid 10001 --create-home --shell /usr/sbin/nologin app
COPY --chown=10001:10001 . /app
USER 10001:10001
There is a reason the USER must be a numeric UID rather than a name. When the image's USER is a name, Kubernetes'
runAsNonRoot cannot tell whether it is root or not, and it
rejects the Pod.
Error: container has runAsNonRoot and image has non-numeric user (app),
cannot verify user is non-root
An unprivileged user cannot open ports below 1024. This is where you may be tempted to add the NET_BIND_SERVICE
capability, but it is usually much simpler to change the app port to 8080 and expose it as 80 at the Service
layer. Leave the choice of adding privileges for last.
Capabilities split root's privileges into about 40 pieces, and the container default set contains
14 of them (CapEff: 00000000a80425fb). It includes cap_chown, cap_dac_override,
cap_fowner, cap_setuid, cap_setgid, cap_net_bind_service, cap_net_raw,
cap_sys_chroot, cap_mknod, cap_setfcap, and others, while SYS_ADMIN, NET_ADMIN, and
SYS_PTRACE are left out.
The recommended combination looks like this.
--cap-drop=ALL --cap-add=NET_BIND_SERVICE
--security-opt no-new-privileges:true
--read-only --tmpfs /tmp
--user 10001:10001
The default seccomp profile blocks about 44 of the 400-odd system calls. Among them are
open_by_handle_at, which was actually used in a container escape, as well as keyctl and kexec_load.
What it looks like in the field
--privileged is a switch that turns off all of the protections above at once. And mounting the Docker socket amounts to the same thing. With that socket, you can start a new privileged container with the host
root filesystem mounted. CI runner
configurations often leave these two on out of habit, and the moment they do, the rest of the security settings become decoration.
Real incidents follow this same pattern. CVE-2019-5736 overwrote the host's runc binary through /proc/self/exe, and CVE-2024-21626 let an attacker reach the host filesystem with nothing more than a manipulated WORKDIR, because runc did not close a file descriptor.
Both were runtime bugs, not kernel bugs, but they worked because the container sits
under the same kernel as the host.
How far can you actually cut privileges
"Not running as root" is only the start. What you can cut back in a container comes in four layers: user, privileges (capabilities), filesystem, and system calls. If you touch all four layers, what an attacker can do after a compromise shrinks greatly.
securityContext:
runAsNonRoot: true
runAsUser: 10001
allowPrivilegeEscalation: false # setuid 로 권한을 되찾는 길을 막는다
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
seccompProfile:
type: RuntimeDefault
allowPrivilegeEscalation: false quietly does a lot of work. If it is left on, a setuid binary inside the container can raise privileges again. If you lower the user but do not turn this off, you have done only half the job.
Drop all capabilities and add back only what you need. If you must open a port below 1024,
add just NET_BIND_SERVICE. A better answer is to listen on 8080 from the start and have the Service
map it to 80, which removes any reason to add privileges at all.
A read-only root reveals where writes are needed. When you turn it on, things usually break at first,
and the places where they break are where the program writes. If you attach an emptyDir only at /tmp and the cache path, nothing else can be written. The real value of this setting is that the attacker has nowhere to download and store tools.
The most common stumbling block is file ownership. If the image was built with root ownership and you run it as a different user, it cannot read the files. It is better to match them in advance
with COPY --chown=10001:10001 when building the image. For volumes, fsGroup changes the group
at mount time.
The default seccomp profile is often enough on its own. RuntimeDefault blocks dozens of dangerous system calls such as
unshare and ptrace. A custom profile is more precise, but it carries a maintenance cost: it breaks silently when the kernel or runtime changes.
Turning on the default alone gets you most of the benefit.
What you will do in the next lab
You confirm that the default image starts as root, create unprivileged execution in two ways, with a runtime flag and with a Dockerfile, and then apply a read-only root, removal of all capabilities, and blocking of privilege escalation one at a time. At the end, you build an image that meets all of these conditions yourself.