CKS — Kubernetes Security Specialist
Knowing What Comes In Comes First
In one line
The core question of supply chain security is not "is this image safe" but "what exactly are the bytes running in the cluster right now, where did they come from, and have they been swapped out?"
Why this was needed
A tag like myapp:v1.2 is a mutable pointer. A different image can be pushed to the same tag,
and latest even more so. If the Pod spec has only a tag, no one can prove what bytes are running on the node right now.
Even if you roll back, the same tag may point to something different.
A digest (@sha256:...) is a hash of the content, so it eliminates this problem at the root.
The same logic applies to language packages. In one project, there were 37 directly declared dependencies, but 1,123 packages were actually installed, and 486 of them were deployed into the production runtime. 37 were reviewed and chosen; 1,123 were trusted. The integrity hash in a lock file guarantees not "this package is safe" but only "this package is the same as the last one I saw." The difference between those two sentences is almost all of supply chain security.
How it works
Controls are stacked in four layers.
| Layer | What it answers | Tools and fields |
|---|---|---|
| Pinning | What are the bytes running right now | Image digest, imagePullPolicy |
| Origin | Where did it come from | Allowed registry list, imagePullSecrets |
| Attestation | Was it built by our pipeline | Signature (cosign), SBOM, provenance attestation |
| Enforcement | How do you block what isn't | Admission policy (VAP, Kyverno, Gatekeeper) |
imagePullPolicy is often misunderstood. Always checks the manifest against the registry every time,
and IfNotPresent pulls only when it isn't on the node. With tags, Always looks safe,
but it also means you face the problem of tags being able to change every single time. If you pin by digest,
the content by definition doesn't change, so it is reasonable to leave IfNotPresent, and a registry outage
doesn't spread into a deployment outage.
A misconception about scanners also needs to be addressed. An image scanner does only two things. It unpacks the image layers
and builds a list of installed packages, and then matches that list against a vulnerability database. So
binaries fetched with curl and copied in, or libraries built directly from source, aren't visible at all.
This is the first reason a clean scan result doesn't mean it is safe.
The last layer is admission. A scan result or a signature has no enforcement power on its own. What keeps an image pushed by hand around the pipeline from entering the cluster is only an admission policy. Until then, a scan gate is closer to a bypassable recommendation.
What it looks like in the field
When you first attach a scan report, you usually see numbers like this. A single payments:1.4.2 image has
1,247 findings (LOW 812, MEDIUM 289, HIGH 137, CRITICAL 9). If you toss that into the team channel as is, nothing
happens. Add a single --ignore-unfixed and it drops to 4 findings (HIGH 3, CRITICAL 1).
Nothing became safer, but only the items you can act on right now remain. If you then cross-check with the CISA KEV list,
exactly 1 was confirmed to be actually exploited. Today's work is that one.
And what eliminated the remaining 1,200 or so was not individual patches but replacing the base image.
Moving the same application from node:22 (432 packages) to a distroless base (19) took
the fixable CRITICAL findings from 9 to 0. The disappearance of the shell and the package manager is
not a bonus but the essence, because it means the image has no tools for an intruder to use.
In exchange, you can no longer get in with kubectl exec, so debugging changes to attaching an ephemeral container.
In an air-gapped network, registry control is also an availability issue. containerd 1.x read the pause image
address from the sandbox_image key, but in 2.x it moved to the sandbox key under pinned_images.
When a cluster that had switched to an internal registry is upgraded to 2.x, the old key is silently ignored
and it falls back to the default, the public registry. The result is that not a single Pod comes up on that node.
Because the whole node dies rather than a specific workload, it is easy to mistake for a network outage.
Three layers that control images
The supply chain domain comes up in the form of "stop this image from being used." There are three places to block it, and which layer is being asked about differs by question.
First, the container runtime. You use node configuration to allow only certain registries or make it verify signatures. It applies across the whole cluster, but you must configure it on every node.
Second, admission control. You block it with a policy inside the cluster. It is the place that comes up
most often on the exam. ImagePolicyWebhook is a dedicated mechanism that asks an external service
whether an image is allowed, and policy engines (Kyverno, Gatekeeper) set more general conditions.
# ImagePolicyWebhook 을 켤 때 함께 필요한 것
# --admission-control-config-file 로 설정 파일을 주고,
# 그 파일이 kubeconfig 를 가리키고, 둘 다 정적 파드에 마운트되어야 한다.
With defaultAllow: false, everything is rejected when the webhook can't be reached. It is safe, but
if the webhook dies, nothing can be deployed to the cluster. Exam questions usually ask about this value.
Third, the image itself. Signatures and SBOMs. Signature verification answers "who built it," and a scan answers "what is in it." The two don't substitute for each other — there can be a signed vulnerable image and a clean anonymous image.
Pin by digest, not by tag. Questions about blocking :latest with a policy
come up often. A tag moves, so what you verified and what runs can diverge.
kubectl get pods -A -o jsonpath='{range .items[*]}{.spec.containers[*].image}{"
"}{end}' | sort -u | grep -vE '@sha256:'
Static analysis also belongs to the supply chain. Tools that check manifests before deployment (kubesec, kube-linter, trivy config) come up as questions. The point is that blocking it before it goes to the cluster is the cheapest layer.
What you will do in the next lab
First you pin an image by digest, build an allowed registry list, and attach a Secret for private registry authentication to a ServiceAccount. You write the signature and SBOM verification policy as a manifest file. In the following lab, you build yourself a ValidatingAdmissionPolicy that can actually be applied to the cluster, which rejects "images that use only a tag" with a CEL expression.