CKA — Kubernetes Administrator
Controlling Scheduling
Goal
You use all seven means of deciding which node a Pod runs on, and at the end you design placement that takes into account the cluster state you created in the earlier steps (a cordoned node and a tainted node).
Why it matters
Scheduling carries little weight on the CKA by itself, but half of the troubleshooting problems are actually scheduling problems. Looking at a Pending Pod, you need to be able to read which condition in the filter phase it failed on.
topologySpreadConstraints in particular is widely misunderstood. The candidate domains used when computing skew are the nodes that pass the Pod's nodeAffinity/nodeSelector. Taints are ignored by default when counting. So a node the Pod cannot use remains as a domain with 0 Pods and widens the skew, and with DoNotSchedule it is common for all the remaining Pods to end up Pending. Knowing that node affinity is the key to narrowing the candidates is the core of this lab.
Steps
- Create the namespace
cka-sched. Labellab-node-0withtopology.kubernetes.io/zone=zone-a,lab-node-1withzone-b, andlab-node-2withzone-c, and additionally labellab-node-0withdisktype=ssd. Then create a Podssd-pod(imagenginx:1.27) with the nodeSelectordisktype=ssdso that it runs onlab-node-0. - Label
lab-node-2withdisktype=hdd, and place the Podaffinity-requiredusing only a required node affinity (disktype In [hdd]) so that it lands onlab-node-2. Do not use nodeSelector. - Create the Deployment
pref(2 replicas, imagenginx:1.27) and attach a preferred node affinity. The weight is 50 and the condition isdisktype In [ssd]. - Create the Deployment
spread(3 replicas, Pod labelapp=spread) and attach a required Pod anti-affinity. The labelSelector isapp=spreadand the topologyKey istopology.kubernetes.io/zone. - Create the Deployment
drain-demo(2 replicas, imagenginx:1.27), then cordon and drainlab-node-1to empty it completely.drain-demomust come back at 2/2 on other nodes. - Taint
lab-node-2withworkload=batch:NoSchedule. Give the Podbatch-poda toleration for that taint and the nodeSelectordisktype=hdd, and place it onlab-node-2. - Create the PriorityClass
cka-high. The value is 100000, globalDefault is false, and preemptionPolicy isPreemptLowerPriority. Assign this class to the Podimportant. - Create the Deployment
evenly(4 replicas, Pod labelapp=evenly). The topologySpreadConstraints are maxSkew 1, topologyKeytopology.kubernetes.io/zone, whenUnsatisfiableDoNotSchedule, and labelSelectorapp=evenly. Also attach a required node affinitydisktype In [ssd, hdd]and a toleration forworkload=batch:NoSchedule, so that the 4 Pods are placed 2:2 across two zones.
Reference
kubectl get pods -n cka-sched -o wideshows right away which node each Pod landed on.- A drain usually goes all the way through only when you add
--ignore-daemonsets --delete-emptydir-data --force. - Common mistake 1: stopping at cordon in step 5. A cordon blocks only new Pods.
- Common mistake 2: leaving out the node affinity in step 8. Then the cordoned
lab-node-1remains as a zone with 0 Pods, the skew becomes 2, and half of the Pods stay Pending.
Node labels and nodeSelector
Create the namespace cka-sched. Label lab-node-0 with topology.kubernetes.io/zone=zone-a, lab-node-1 with zone-b, and lab-node-2 with zone-c, and additionally label lab-node-0 with disktype=ssd. Then create a Pod ssd-pod (image nginx:1.27) with the nodeSelector disktype=ssd so that it runs on lab-node-0.
nodeSelector is a map at the top level of the Pod spec. The label must match exactly, and you cannot use expressions.
Required node affinity
Label lab-node-2 with disktype=hdd, and place the Pod affinity-required using only a required node affinity (disktype In [hdd]) so that it lands on lab-node-2. Do not use nodeSelector.
Under affinity.nodeAffinity, requiredDuringSchedulingIgnoredDuringExecution holds a nodeSelectorTerms array. The operator of matchExpressions is In. Do not use nodeSelector.
Preferred node affinity
Create the Deployment pref (2 replicas, image nginx:1.27) and attach a preferred node affinity. The weight is 50 and the condition is disktype In [ssd].
The preferred side is an array, and each item has a weight and a preference. The Pod is placed even if it cannot be satisfied, so where it landed is not graded.
Spread Pods with Pod anti-affinity
Create the Deployment spread (3 replicas, Pod label app=spread) and attach a required Pod anti-affinity. The labelSelector is app=spread and the topologyKey is topology.kubernetes.io/zone.
Each item of podAntiAffinity has a labelSelector and a topologyKey. Each node has a different zone label, so spreading by zone ends up separating the nodes.
Empty a node
Create the Deployment drain-demo (2 replicas, image nginx:1.27), then cordon and drain lab-node-1 to empty it completely. drain-demo must come back at 2/2 on other nodes.
A cordon blocks only new Pods and leaves the Pods that are already running. To actually empty the node you need a drain, and it proceeds only when you add the options related to DaemonSets and emptyDir.
Taints and tolerations
Taint lab-node-2 with workload=batch:NoSchedule. Give the Pod batch-pod a toleration for that taint and the nodeSelector disktype=hdd, and place it on lab-node-2.
A taint has the form key=value:effect. A toleration must set the operator to Equal and match the value as well. Which node it goes to must be specified separately.
Attach a PriorityClass
Create the PriorityClass cka-high. The value is 100000, globalDefault is false, and preemptionPolicy is PreemptLowerPriority. Assign this class to the Pod important.
A PriorityClass is cluster-scoped. Be careful: setting globalDefault to true affects the Pods of the entire cluster. If you write only the name in the Pod, spec.priority is filled in automatically.
Putting it together: spread evenly over the remaining nodes
Create the Deployment evenly (4 replicas, Pod label app=evenly). The topologySpreadConstraints are maxSkew 1, topologyKey topology.kubernetes.io/zone, whenUnsatisfiable DoNotSchedule, and labelSelector app=evenly. Also attach a required node affinity disktype In [ssd, hdd] and a toleration for workload=batch:NoSchedule, so that the 4 Pods are placed 2:2 across two zones.
First count how many nodes in the cluster right now can accept Pods. A cordoned node must be removed from the candidates, and a tainted node can be used only with a toleration. The key to reducing the candidates is node affinity.