CKS — Kubernetes Security Specialist
Tightening With seccomp, AppArmor and Capabilities
Goal
You get into your hands how to reduce what a container can request from the kernel using only the Pod spec, and you build a detection procedure that finds privileged containers remaining in the cluster.
Why it matters
Containers share the host kernel. Escape vulnerabilities mostly strike the kernel through system calls or capabilities.
So defense is designed not to block the intrusion but to reduce in advance what an intruding process
can do. seccomp restricts the callable system calls, AppArmor the accessible
files and capabilities, and capabilities the pieces of root privilege. Because the three layers overlap,
even if one is breached, the others remain. Conversely, privileged: true nullifies all three layers at once,
so counting privileged containers in a production cluster is itself a measure of risk.
Note: this lab environment has no real container runtime. You can't confirm that seccomp and AppArmor profiles are really enforced, and it doesn't check whether the profile files exist on the node. Grading looks only at whether the Pod spec's fields are written exactly. What the CKS exam grades is, in the end, this spec too.
Steps
-
Create the namespace
cks-sysand create a Podseccomp-default(imagenginx:1.27-alpine) in it. The Pod-levelspec.securityContext.seccompProfile.typeisRuntimeDefault. -
Create a Pod
seccomp-customincks-sys. Thetypeof the Pod-level seccompProfile isLocalhost, andlocalhostProfileisprofiles/audit.json. -
Create a Pod
apparmor-app(container nameapp, imagenginx:1.27-alpine) incks-sysand apply the AppArmor profilek8s-custom-profile. Thetypeof the container'ssecurityContext.appArmorProfileisLocalhost, andlocalhostProfileisk8s-custom-profile. (Using the old-formatcontainer.apparmor.security.beta.kubernetes.io/app: localhost/k8s-custom-profileannotation also passes.) -
Create a Pod
capped(container nameapp) incks-sys. In the container securityContext,capabilities.dropis the single entryALL,capabilities.addis the single entryNET_BIND_SERVICE, andallowPrivilegeEscalationisfalse. -
Create a Pod
no-host-nsincks-sys. Explicitly sethostNetwork,hostPID,hostIPC, andshareProcessNamespaceall tofalse.There is a fact hidden here that you must know for audits. In the PodSpec, the first three fields are
bool+omitempty, so even if you writefalse, it is dropped entirely at the serialization stage. That is, from the stored object alone you can't tell "I wrote false" from "I didn't write it at all." So the actual audit criterion is not "is it written as false" but "is it not true." On the other hand,shareProcessNamespaceis a*bool, sofalseis kept as is — the grader also reflects this difference as it is. After applying it, check for yourself withkubectl get pod no-host-ns -n cks-sys -o yaml. -
Create a Pod
legacy-agent(container nameagent) incks-sys, but set the container securityContext'sprivilegedtotrue(reproducing a legacy workload). Then write a detection script in/root/cks-system-hardening/find-privileged.shand save its output to/root/cks-system-hardening/privileged.txt. The output is the Pods in thecks-sysnamespace that have privileged containers (including init containers), in the formatcks-sys/<파드이름>(the placeholder is the Pod name), one per line, sorted alphabetically. -
Create a Pod
hardened(container nameapp) incks-sys. In the container securityContext put all ofreadOnlyRootFilesystem: true,runAsNonRoot: true,runAsUser: 1000,allowPrivilegeEscalation: false, andcapabilities.drop: [ALL], and set the Pod-levelseccompProfile.typetoRuntimeDefault. Mount/tmp, which needs writing, as an emptyDir volume namedtmp.
Notes
- It is faster to pull out a draft with
kubectl run seccomp-default --image=nginx:1.27-alpine -n cks-sys --dry-run=client -o yaml > pod.yamland edit it. - Check the field names with
kubectl explain pod.spec.securityContext.seccompProfile. - Detection script hint: with
kubectl get pods -n cks-sys -o json | jq -r '...', scan both.spec.containers[]and.spec.initContainers[]. - Common mistake 1: putting
capabilitiesin the Pod-level securityContext. capabilities exist only at the container level. - Common mistake 2: writing an absolute path in
localhostProfile. It must be a path relative to the node's seccomp root.
The runtime default seccomp profile
Create the namespace cks-sys and create a Pod seccomp-default (image
nginx:1.27-alpine) in it. The Pod-level spec.securityContext.seccompProfile.type is
RuntimeDefault.
seccompProfile can be written in the Pod-level spec.securityContext as well as at the container level. Here it is the Pod level.
Specify a custom seccomp profile
Create a Pod seccomp-custom in cks-sys. The type of the Pod-level seccompProfile is
Localhost, and localhostProfile is profiles/audit.json.
If the type is Localhost, localhostProfile must be present with it, and the path is relative to the node's seccomp root.
Attach an AppArmor profile
Create a Pod apparmor-app (container name app, image nginx:1.27-alpine) in cks-sys and
apply the AppArmor profile k8s-custom-profile. The type of the container's
securityContext.appArmorProfile is Localhost, and localhostProfile is
k8s-custom-profile. (Using the old-format
container.apparmor.security.beta.kubernetes.io/app: localhost/k8s-custom-profile
annotation also passes.)
From 1.30 you use the container securityContext.appArmorProfile field. On an older cluster it is the container.apparmor.security.beta.kubernetes.io/<컨테이너이름> annotation format (the placeholder is the container name). Either one will do.
Drop all capabilities and get back just one
Create a Pod capped (container name app) in cks-sys. In the container securityContext,
capabilities.drop is the single entry ALL, capabilities.add is the single entry NET_BIND_SERVICE, and
allowPrivilegeEscalation is false.
Put ALL in drop and list in add only what is truly needed. Don't attach the CAP_ prefix to capability names.
Block host namespaces
Create a Pod no-host-ns in cks-sys. Explicitly set hostNetwork, hostPID, hostIPC, and
shareProcessNamespace all to false.
hostNetwork, hostPID, hostIPC, and shareProcessNamespace are all top-level fields of the Pod spec. Write all four as false, but if you read the stored object back, the first three disappear and only shareProcessNamespace remains — see item 5 of the instructions for why.
Detect privileged containers
Create a Pod legacy-agent (container name agent) in cks-sys, but set the container
securityContext's privileged to true (reproducing a legacy workload).
Then write a detection script in /root/cks-system-hardening/find-privileged.sh and
save its output to /root/cks-system-hardening/privileged.txt. The output is the Pods in the cks-sys
namespace that have privileged containers (including init containers), in the format
cks-sys/<파드이름> (the placeholder is the Pod name), one per line, sorted alphabetically.
Scan kubectl get pods -o json with jq, and look at initContainers as well so you don't miss any. The output format is 네임스페이스/파드이름 (the placeholders are the namespace and the Pod name).
Putting together a read-only root filesystem
Create a Pod hardened (container name app) in cks-sys. In the container securityContext put all of
readOnlyRootFilesystem: true, runAsNonRoot: true, runAsUser: 1000,
allowPrivilegeEscalation: false, and capabilities.drop: [ALL],
and set the Pod-level seccompProfile.type to RuntimeDefault. Mount /tmp, which needs writing, as
an emptyDir volume named tmp.
Making the root read-only blocks temporary file paths. Mount an emptyDir volume at that path. You gather all the fields you wrote in the earlier steps into one Pod.