CKS — Kubernetes Security Specialist
Different Blocking Points Mean Different Places to Fix
In one line
Practicing writing a configuration is different from practicing confirming that the configuration is enforced. Accidents always happen in the latter.
Why do it where it is enforced
In the earlier modules, you used securityContext and Pod Security many times. But the places where those labs ran had neither a kubelet nor a container runtime, so even violating Pods just became Running.
You practiced writing configuration, but you did not practice confirming that the configuration is enforced. And in practice, the place where accidents happen is the latter — because a state where all the configuration is written but not applied looks normal.
The same "it doesn't work" is scattered across three layers
| Layer | What lives here | Symptom | When you see it |
|---|---|---|---|
| Admission | Pod Security Admission, policy engines | kubectl apply is rejected |
Before deployment (CI) |
| kubelet | runAsNonRoot, the image USER check |
The Pod is created, CreateContainerConfigError |
Right after deployment |
| Runtime | Read-only root, capabilities, seccomp | The Pod is Running, only the app fails |
When it takes that code path |
The best is admission. Feedback comes immediately, and wrong things never enter the cluster at all.
The worst is runtime. The Pod is running fine, so the symptom looks like an application error. It is also common to be fine normally and fail only on a specific code path.
So there is an order. Don't leave to a later layer what can be blocked at an earlier layer.
Start Pod Security Admission with warn
If you apply enforce right away, workloads that are already running are all blocked at the next redeployment. That is itself an outage.
pod-security.kubernetes.io/warn: restricted # 먼저 이것만
pod-security.kubernetes.io/enforce: restricted # 경고가 0 이 된 뒤에
The order is not to cut first and then count, but to count first and then cut.
What breaks when you tighten a container
If you know what actually fails when you turn on security settings, you can fix the image side before you experience an outage. Here are just the ones you run into often.
A read-only root filesystem. Most apps write temporary files somewhere. Logs, caches, sockets, and the temporary directories used by language runtimes. If you make the root read-only, all of these get blocked, and the symptom usually shows up as a permission error in the application, so it takes time to connect it to the security setting. The fix is to attach an emptyDir to every path that needs writing, and rather than guessing where it writes, it is faster to bring it up once in the tightened state and collect the failing paths.
Running as a non-root user. The first thing you hit is that you can't open ports below 1024. An image that listened on port 80 inside the container has to change to 8080, and the Service's targetPort has to be fixed along with it. And if file owners inside the image are root, the new user may not be able to read them, so you have to hand over ownership when building the image.
Dropping all capabilities. Most apps need nothing, but if you must open a low port you need NET_BIND_SERVICE, and if you use ping you need NET_RAW. You have to go in the order of adding back only what is needed; if you leave everything in place without knowing what is needed, tightening is meaningless.
The seccomp default profile. System calls that aren't allowed come back as EPERM, and if a library swallows that, the symptom shows up in an unexpected way. The reason the container labs in this lab environment are blocked is of the same family.
To sum up, the order is this. Bring it up first with the tightened configuration and collect what fails, open exceptions only for what failed, and then work on removing that exception list from the image side. Conversely, if you start with "loosen everything for now and tighten later," that later never comes.
What really matters in practice
Don't leave to a later layer what can be blocked at an earlier layer. If it is caught at admission, CI knows immediately, but if it is caught at runtime, the Pod keeps running fine and fails only on a specific code path, looking like an application error. Even for the same rule, where you enforce it changes the response cost by a factor of ten.
Apply enforce after the warnings reach 0. If you apply it right away, workloads that are already running are all blocked at the next redeployment, and that is itself an outage. You must keep the order of counting with warn first and then cutting.
A runAsNonRoot failure doesn't look like a Pod creation failure. The Pod is created and only CreateContainerConfigError remains, so it takes time to suspect the image's USER. Specifying USER explicitly when building the image is the cheapest prevention.
In the next lab, you check these three one at a time on a real cluster.