CKS — Kubernetes Security Specialist
A Container Is a Restriction, Not an Isolation
In one line
Containers share the host kernel. So system hardening is less about "strengthening isolation" and more about "reducing what can be requested from the shared kernel."
Why this was needed
A virtual machine has its own kernel. Even if you attack a kernel vulnerability from the guest, it is hard to cross the hypervisor boundary. Containers are different. Every Pod on a node uses the same kernel, and the kernel's system call interface has hundreds of calls. Container escape CVEs mostly strike somewhere on this surface.
So the defensive logic changes. Instead of preventing the intrusion itself, you reduce in advance the actions that an intruding process can request from the kernel. Linux has mechanisms for this purpose in three layers, and Kubernetes exposes all three as Pod spec fields.
| Mechanism | What it restricts | Pod spec field |
|---|---|---|
| seccomp | The list of callable system calls | securityContext.seccompProfile |
| AppArmor / SELinux | Accessible files, network, and capabilities (MAC) | securityContext.appArmorProfile |
| capabilities | The fine-grained pieces of root privilege | securityContext.capabilities |
How it works
For seccomp, the two types RuntimeDefault and Localhost are all there is in practice. RuntimeDefault uses
the default block list held by the container runtime, and Localhost points to a JSON profile under the node's
<seccomp-root>/profiles through the localhostProfile path.
If the profile file isn't on the node, the Pod can't start. Writing it in the spec and having it on the node are
separate things.
AppArmor was promoted from an annotation to a proper field in Kubernetes 1.30. The old form was the
container.apparmor.security.beta.kubernetes.io/<컨테이너이름> (the placeholder is the container name) annotation, and now it is
securityContext.appArmorProfile.type and localhostProfile. On the exam, both can come up depending on the cluster
version, so you need to know both forms. Here too the profile must be loaded on the node
in advance, and you check it with aa-status.
Capabilities are the easiest to understand and have the biggest effect. A container's root is not real root but a UID 0 with
a bundle of capabilities. If you discard everything with drop: ["ALL"] and then add only what is really needed,
a web server that needs to open port 80, for example, needs just NET_BIND_SERVICE.
privileged: true nullifies all of this. It grants every capability, opens access to every device,
and even disables the AppArmor, SELinux, and seccomp profiles. That is why finding privileged containers
is a staple CKS question.
Sharing host namespaces is the same layer. With hostPID: true, every process on the node is visible from the container,
and you can read another process's environment variables through /proc/<PID>/environ. If a workload that injected its secrets
as environment variables is next to it, that is the end. hostNetwork makes network policy
effectively meaningless, and hostIPC erases the shared memory boundary.
Last is readOnlyRootFilesystem: true. It erases the path by which an intruder drops a binary or swaps out an existing
binary. Paths that need writing are mounted separately as an emptyDir.
What it looks like in the field
Reading container image scan results makes the need for this layer clear. A single node:22 image has 432 packages
installed and produces a vulnerability report of 1,247 findings (CRITICAL 9, HIGH 137).
Yet if you count with ldd the shared libraries the application actually links against, there are 8.
The rest came along with the base image, and most of it is shells, package managers, and utilities.
For an intruder, that is a toolbox. Deleting them from the image is best, and if you can't delete them,
reducing what those tools can request from the kernel with capabilities and seccomp is the next best thing.
The same story repeats at the CRI layer. containerd translates the Pod spec's SecurityContext into the OCI spec.
readOnlyRootFilesystem: true becomes root.readonly = true,
capabilities goes down to process.capabilities, and seccompProfile to linux.seccomp.
And privileged: true is expanded during this translation into granting every capability, allowing every device,
and disabling AppArmor, SELinux, and seccomp. It is a structure in which a single field brings down three layers
at once. Whether it was actually applied is confirmed by reading the final spec on the node with crictl inspect.
What you will do in the next lab
This environment has no real container runtime, so you can't confirm that profiles are really enforced. Instead, you practice what the CKS actually grades, namely writing the Pod spec's fields exactly. You build by hand the two seccomp types, an AppArmor profile, capability drop/add, blocking host namespaces, and even a detection script that finds privileged containers.