TT Lab
Get started
Learn Learning paths Courses

Container Security

Tightening a Container to Least Privilege

Continue in TT Lab

This lab runs on a real VM

This box is not a Pod but a virtual machine started by KubeVirt. A Linux kernel runs separately, systemd actually manages services, and docker is not an imitation but a real Docker engine. A container started with docker run actually becomes a process, and docker exec and docker logs work as expected.

This lab used to run inside a Pod. The box had dropped all kernel privileges, so the step that starts a container was blocked, and the lab worked around that by unpacking the image archive directly. The workaround is no longer needed.

There are two things to know.

Goal

Build both ways of running a container as an unprivileged user (a runtime flag and building it into the image), and apply a read-only root filesystem, removal of all capabilities, and blocking of privilege escalation one at a time. At the end, you complete an image that meets these conditions yourself.

Why it matters

The reason the claim "containers are isolated, so it is fine to be root inside" is wrong is simple. If you do not use user namespaces, UID 0 in the container is the same value as UID 0 on the host, and on any path opened up by a volume, that privilege applies as it is. So least privilege should be the default, not an option. The habit of writing USER as a number matters especially. Kubernetes' runAsNonRoot cannot tell that a user given by name is not root, so it rejects the Pod, and the first time you see that error message it takes a long time to find the cause. The same goes for capabilities. Removing only what you do not need always leaves something out eventually, but dropping everything and then adding back only the one you need leaves nothing out.

Steps

  1. Create the /root/sec1 directory, and save only the numeric UID of running alpine:3.20 with no user specified to /root/sec1/default-uid.txt. The file must contain only 0. You can answer this with certainty without starting a container. The user that runs is recorded in the image config's config.User field, and if it is empty, the runtime uses uid 0. Check with skopeo inspect --config oci-archive:/opt/images/alpine_3.20.tar | jq -r '.config.User'.
  2. Run a sec-u container from alpine:3.20 in the background with --user 1000:1000, and keep it running until grading time.
  3. Write /root/sec1/Dockerfile and build the labhub/sec:v1 image. Specify numeric UID 10001 with the USER directive so that docker run --rm labhub/sec:v1 id -u prints 10001.
  4. As an unprivileged user, try to create a file under /etc, and save the resulting Permission denied error, including standard error, to /root/sec1/denied.txt.
  5. Run a sec-ro container in the background with --read-only, try to write a file inside it, and save the resulting Read-only file system error to /root/sec1/ro.txt.
  6. Run a sec-caps container in the background with --cap-drop=ALL --cap-add=NET_BIND_SERVICE.
  7. Run a sec-nnp container in the background with --security-opt no-new-privileges:true.
  8. Write /root/sec1/hardened.Dockerfile and build the labhub/sec:v2 image. It must meet all four of the following.
    • FROM is pinned to a specific tag and is not latest (for example, alpine:3.20)
    • The last USER is a numeric UID and the value is 10000 or higher
    • The org.opencontainers.image.source label is present
    • When actually run, id -u prints that UID

Notes

The default image starts as root

Create the /root/sec1 directory, and save only the numeric UID of running alpine:3.20 with no user specified to /root/sec1/default-uid.txt. The file must contain only 0. You can answer this with certainty without starting a container. The user that runs is recorded in the image config's config.User field, and if it is empty, the runtime uses uid 0. Check with skopeo inspect --config oci-archive:/opt/images/alpine_3.20.tar | jq -r '.config.User'.

Save only one numeric UID, not the full output of id. If you put in the uid=0(root) gid=0... form as it is, several numbers are mixed in and the check fails. There is an option that prints only the number.

Change the UID with a runtime flag

Run a sec-u container from alpine:3.20 in the background with --user 1000:1000, and keep it running until grading time.

There is a flag that changes the user at run time without rebuilding the image. Give it in UID:GID form. Grading runs commands inside the container, so the container must keep running.

Make the image itself non-root

Write /root/sec1/Dockerfile and build the labhub/sec:v1 image. Specify numeric UID 10001 with the USER directive so that docker run --rm labhub/sec:v1 id -u prints 10001.

To stay safe even when the person running it forgets the flag, the default must be baked into the image. Create the user first, then switch with USER, and specify it by number, not by name. alpine has adduser instead of useradd.

What happens when you have no privileges

As an unprivileged user, try to create a file under /etc, and save the resulting Permission denied error, including standard error, to /root/sec1/denied.txt.

The error message comes out on standard error, not standard output. If you do not capture standard error when redirecting, the file is empty and the check fails.

Read-only root filesystem

Run a sec-ro container in the background with --read-only, try to write a file inside it, and save the resulting Read-only file system error to /root/sec1/ro.txt.

Start the container with the flag that makes the root filesystem read-only, try to write inside it, and save the resulting error. Here too you must capture standard error. In a real service, you would mount tmpfs separately on the paths that need writes.

Drop every capability and add back just one

Run a sec-caps container in the background with --cap-drop=ALL --cap-add=NET_BIND_SERVICE.

The order matters. It must not be the approach of removing only what you do not need, but dropping everything and then adding back only the one you need. The only one to add is the capability used for binding low ports.

Block the privilege escalation path

Run a sec-nnp container in the background with --security-opt no-new-privileges:true.

There is a security option that blocks, at the kernel level, the path by which running a binary with the setuid bit raises privileges. Specify it with --security-opt and add :true after the value.

Complete the hardened image

Write /root/sec1/hardened.Dockerfile and build the labhub/sec:v2 image. It must meet all four of the following.

You must meet all four at the same time — a pinned tag (no latest), a numeric UID of 10000 or higher, a source label, and the container actually running as that UID. The label key is an OCI standard name, so the spelling must be exact.