Tightening a Container to Least Privilege
This lab runs on a real VM
This box is not a Pod but a virtual machine started by KubeVirt. A Linux kernel
runs separately, systemd actually manages services, and docker is not an imitation but
a real Docker engine. A container started with docker run actually becomes a process,
and docker exec and docker logs work as expected.
This lab used to run inside a Pod. The box had dropped all kernel privileges, so the step that starts a container was blocked, and the lab worked around that by unpacking the image archive directly. The workaround is no longer needed.
There are two things to know.
- The first start takes a little over a minute. This is because the VM boots and installs Docker. It is slower than a Pod lab (usually 40 seconds).
- There is no browser preview. Only one grading port is open for connections into the VM. If you started a web server, check it with
curlfrom inside the VM.
Goal
Build both ways of running a container as an unprivileged user (a runtime flag and building it into the image), and apply a read-only root filesystem, removal of all capabilities, and blocking of privilege escalation one at a time. At the end, you complete an image that meets these conditions yourself.
Why it matters
The reason the claim "containers are isolated, so it is fine to be root inside" is wrong is simple.
If you do not use user namespaces, UID 0 in the container is the same value as UID 0 on the host, and on any path opened up by a volume, that privilege applies as it is. So least privilege should be the default, not an option. The habit of writing USER as a number matters especially.
Kubernetes' runAsNonRoot cannot tell that a user given by name is not root, so it
rejects the Pod, and the first time you see that error message it takes a long time to find the cause.
The same goes for capabilities. Removing only what you do not need always leaves something out eventually, but dropping everything
and then adding back only the one you need leaves nothing out.
Steps
- Create the
/root/sec1directory, and save only the numeric UID of runningalpine:3.20with no user specified to/root/sec1/default-uid.txt. The file must contain only0. You can answer this with certainty without starting a container. The user that runs is recorded in the image config'sconfig.Userfield, and if it is empty, the runtime uses uid 0. Check withskopeo inspect --config oci-archive:/opt/images/alpine_3.20.tar | jq -r '.config.User'. - Run a
sec-ucontainer fromalpine:3.20in the background with--user 1000:1000, and keep it running until grading time. - Write
/root/sec1/Dockerfileand build thelabhub/sec:v1image. Specify numeric UID 10001 with theUSERdirective so thatdocker run --rm labhub/sec:v1 id -uprints10001. - As an unprivileged user, try to create a file under
/etc, and save the resultingPermission deniederror, including standard error, to/root/sec1/denied.txt. - Run a
sec-rocontainer in the background with--read-only, try to write a file inside it, and save the resultingRead-only file systemerror to/root/sec1/ro.txt. - Run a
sec-capscontainer in the background with--cap-drop=ALL --cap-add=NET_BIND_SERVICE. - Run a
sec-nnpcontainer in the background with--security-opt no-new-privileges:true. - Write
/root/sec1/hardened.Dockerfileand build thelabhub/sec:v2image. It must meet all four of the following.FROMis pinned to a specific tag and is notlatest(for example,alpine:3.20)- The last
USERis a numeric UID and the value is 10000 or higher - The
org.opencontainers.image.sourcelabel is present - When actually run,
id -uprints that UID
Notes
- To get only the numeric UID, use
id -u. Running plainidprints several numbers mixed together. - To capture standard error in the file as well, use the form
명령 > 파일 2>&1(the placeholders are the command and the file). - Step 4 can be done without a container.
setpriv --reuid=1000 --regid=1000 --clear-groups <명령>(the placeholder is the command to run) drops the uid/gid and runs the command — the same thing a container runtime does with--user. - To create a user on alpine, use the form
adduser -D -u 10001 app. - Write the label like
LABEL org.opencontainers.image.source="https://example.com/repo". - Common mistake 1: if you save the full output of
idin step 1, all the numbers inuid=0(root) gid=0(root)match and the check fails. - Common mistake 2: if you start the containers for steps 2, 5, 6, and 7 with
--rmor with a command that ends soon, they are gone by grading time. - Common mistake 3: writing a name such as
USER appin step 8 fails. It is the same reason Kubernetes rejects it.
The default image starts as root
Create the /root/sec1 directory, and save only the numeric UID of running alpine:3.20 with no user specified to /root/sec1/default-uid.txt. The file must contain only 0.
You can answer this with certainty without starting a container. The user that runs is recorded in the image config's config.User field, and if it is empty, the runtime uses uid 0. Check with skopeo inspect --config oci-archive:/opt/images/alpine_3.20.tar | jq -r '.config.User'.
Save only one numeric UID, not the full output of id. If you put in the uid=0(root) gid=0... form as it is, several numbers are mixed in and the check fails. There is an option that prints only the number.
Change the UID with a runtime flag
Run a sec-u container from alpine:3.20 in the background with --user 1000:1000, and keep it running until grading time.
There is a flag that changes the user at run time without rebuilding the image. Give it in UID:GID form. Grading runs commands inside the container, so the container must keep running.
Make the image itself non-root
Write /root/sec1/Dockerfile and build the labhub/sec:v1 image. Specify numeric UID 10001 with the USER directive so that docker run --rm labhub/sec:v1 id -u prints 10001.
To stay safe even when the person running it forgets the flag, the default must be baked into the image. Create the user first, then switch with USER, and specify it by number, not by name. alpine has adduser instead of useradd.
What happens when you have no privileges
As an unprivileged user, try to create a file under /etc, and save the resulting Permission denied error, including standard error, to /root/sec1/denied.txt.
The error message comes out on standard error, not standard output. If you do not capture standard error when redirecting, the file is empty and the check fails.
Read-only root filesystem
Run a sec-ro container in the background with --read-only, try to write a file inside it, and save the resulting Read-only file system error to /root/sec1/ro.txt.
Start the container with the flag that makes the root filesystem read-only, try to write inside it, and save the resulting error. Here too you must capture standard error. In a real service, you would mount tmpfs separately on the paths that need writes.
Drop every capability and add back just one
Run a sec-caps container in the background with --cap-drop=ALL --cap-add=NET_BIND_SERVICE.
The order matters. It must not be the approach of removing only what you do not need, but dropping everything and then adding back only the one you need. The only one to add is the capability used for binding low ports.
Block the privilege escalation path
Run a sec-nnp container in the background with --security-opt no-new-privileges:true.
There is a security option that blocks, at the kernel level, the path by which running a binary with the setuid bit raises privileges. Specify it with --security-opt and add :true after the value.
Complete the hardened image
Write /root/sec1/hardened.Dockerfile and build the labhub/sec:v2 image. It must meet all four of the following.
FROMis pinned to a specific tag and is notlatest(for example,alpine:3.20)- The last
USERis a numeric UID and the value is 10000 or higher - The
org.opencontainers.image.sourcelabel is present - When actually run,
id -uprints that UID
You must meet all four at the same time — a pinned tag (no latest), a numeric UID of 10000 or higher, a source label, and the container actually running as that UID. The label key is an OCI standard name, so the spelling must be exact.