CKAD — Kubernetes Application Developer
securityContext, Resources, Quotas and QoS
Goal
You specify in manifests what a Pod can do on the node (securityContext, ServiceAccount) and how much it can use (resources, LimitRange, ResourceQuota), and check how the QoS class is determined as a result.
Why it matters
Containers run as root by default, because the image was built that way. If there is even one container escape vulnerability, that root can lead to root on the node. runAsNonRoot: true refuses to start a container that tries to come up as UID 0, and capabilities.drop: ["ALL"] withdraws every privilege the kernel grants. If a web server has to open port 80, you add back only NET_BIND_SERVICE. This is what the principle of least privilege looks like in concrete form.
Understand resources in two layers. requests is the language of the scheduler. A Pod is placed only if the node's remaining allocatable amount is larger than the sum of requests. It has nothing to do with actual usage. limits is the language of the node. CPU is throttled by a cgroup quota, and memory gets OOMKilled the moment it is exceeded. So if you set the memory limits lower than the actual peak, the Pod quietly restarts over and over.
You must understand LimitRange and ResourceQuota as a set. If you create a Pod without requests/limits in a namespace that has a quota, its creation is rejected outright. If you put only a quota on a team namespace, the incident is that nothing starts, so filling in defaults with a LimitRange is practically mandatory.
Steps
- Create the namespace
ckad-secure, and in it create a ServiceAccountapp-sa. - Create a Pod
sa-pod. Imagenginx:1.27,serviceAccountName: app-sa,automountServiceAccountToken: false. - Create a Pod
nonroot. Imagenginx:1.27, and in the Pod-levelsecurityContextsetrunAsNonRoot: true,runAsUser: 1000,fsGroup: 2000. - Create a Pod
hardened. Imagenginx:1.27, container nameapp, and in the container-levelsecurityContextsetallowPrivilegeEscalation: false,readOnlyRootFilesystem: true,capabilities.drop: ["ALL"],capabilities.add: ["NET_BIND_SERVICE"]. - Create a Deployment
api. 2 replicas, labelapp=api, imagenginx:1.27, containerresources.requestsofcpu: 100m/memory: 128Mi, andresources.limitsofcpu: 500m/memory: 512Mi. - Create a LimitRange
defaults. Withtype: Container, setdefault(limits)cpu: 200m/memory: 256Mi,defaultRequestcpu: 100m/memory: 128Mi,maxcpu: "1"/memory: 1Gi, andmincpu: 50m/memory: 64Mi. - Create a ResourceQuota
team-quota.requests.cpu: "2",requests.memory: 4Gi,limits.cpu: "4",limits.memory: 8Gi,pods: "10". - Create two more Pods.
qos-guaranteedmust have requests and limits that are both equal tocpu: 250m/memory: 256Mi, andqos-burstablegets only requests ofcpu: 100m/memory: 128Mi. Both use imagenginx:1.27.
Notes
- Extract a skeleton with
kubectl create deployment api --image=nginx:1.27 --replicas=2 -n ckad-secure --dry-run=client -o yamland fill inresourcesby hand, or create it and then usekubectl set resources deployment api -n ckad-secure --requests=... --limits=.... - After step 7, you cannot create a Pod without requests/limits in this namespace. If you need to recreate a Pod from an earlier step, check that the LimitRange fills in the defaults.
- Common mistake 1: writing
fsGrouporrunAsNonRootunder the container.fsGroupexists only at the Pod level. - Common mistake 2: writing
capabilitiesat the Pod level. It exists only at the container level. - Check:
kubectl get pod qos-guaranteed -n ckad-secure -o jsonpath='{.status.qosClass}'
Namespace and ServiceAccount
Create the namespace ckad-secure, and in it create a ServiceAccount app-sa.
Create it with kubectl create serviceaccount. Create the namespace first, then create it inside.
Assigning a ServiceAccount to a Pod and turning off automatic token mounting
Create a Pod sa-pod. Image nginx:1.27, serviceAccountName: app-sa, automountServiceAccountToken: false.
The field name is spec.serviceAccountName (serviceAccount is an old alias). The boolean field that turns off automatic token mounting is also at the top level of the Pod spec.
Pod-level securityContext
Create a Pod nonroot. Image nginx:1.27, and in the Pod-level securityContext set runAsNonRoot: true, runAsUser: 1000, fsGroup: 2000.
What goes in spec.securityContext are values that apply to the whole Pod. The field that specifies the group that owns volumes exists only at the Pod level.
Container-level securityContext and capabilities
Create a Pod hardened. Image nginx:1.27, container name app, and in the container-level securityContext set allowPrivilegeEscalation: false, readOnlyRootFilesystem: true, capabilities.drop: ["ALL"], capabilities.add: ["NET_BIND_SERVICE"].
capabilities exists only at the container level. Dropping everything and then adding back only what you need is the principle of least privilege. drop and add are each an array of strings.
Specifying requests and limits in a Deployment
Create a Deployment api. 2 replicas, label app=api, image nginx:1.27, container resources.requests of cpu: 100m / memory: 128Mi, and resources.limits of cpu: 500m / memory: 512Mi.
resources is under the container. requests is the value the scheduler looks at, and limits is the value the runtime enforces. The CPU unit m is 1/1000 of a core.
Setting defaults and bounds with a LimitRange
Create a LimitRange defaults. With type: Container, set default (limits) cpu: 200m / memory: 256Mi, defaultRequest cpu: 100m / memory: 128Mi, max cpu: "1" / memory: 1Gi, and min cpu: 50m / memory: 64Mi.
spec.limits is an array and each element has a type. The key that sets the default limits and the key that sets the default requests have different names, so check with kubectl explain limitrange.spec.limits.
Limiting the namespace total with a ResourceQuota
Create a ResourceQuota team-quota. requests.cpu: "2", requests.memory: 4Gi, limits.cpu: "4", limits.memory: 8Gi, pods: "10".
spec.hard is a map and its key names contain a dot, like requests.cpu and limits.memory. Count limits such as pods are given as string values.
Creating the QoS classes you intend (comprehensive)
Create two more Pods. qos-guaranteed must have requests and limits that are both equal to cpu: 250m / memory: 256Mi, and qos-burstable gets only requests of cpu: 100m / memory: 128Mi. Both use image nginx:1.27.
QoS is not a field you set directly but a result determined by the combination of requests/limits. For the top class, requests and limits must be equal for both CPU and memory in every container. Check it with kubectl get pod -o jsonpath='{.status.qosClass}'.