TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

Watch the Scheduler Pick a Seat

Continue in TT Lab

This lab runs on a cluster that is really short on resources

A real k3s is running inside the VM. Because the node's CPU is actually finite, if you make a large request, there really is not enough room and the Pod stays Pending.

On the fake cluster where the other scheduling labs in the CKA course run, there are several fake nodes and no resource pressure, so whatever you request gets placed. So you could not see what the scheduler actually looks at when it refuses.

It takes about 2 minutes to start up the first time.

Goal

You actually make it get blocked to verify the five mechanisms the scheduler uses to choose a place. And you go on to see a lower-priority Pod being evicted.

Why it matters

With scheduling settings, the problem much more often is "if you get it wrong, nothing gets placed anywhere" than "if you use it well, it lands where you want." And the symptom is always the same — the Pod is Pending.

The causes are spread across several layers.

The scheduler always leaves in the events why it could not place a Pod. Training yourself to read that is the whole of this lab.

Steps

Create everything in the csch namespace. There is one node.

  1. Put the node's capacity and allocatable, and how much is requested right now, in /root/csch/capacity.txt.
  2. Make three Pods, hog-a, hog-b, hog-c, compete for room. The first two must be placed and the third must be Pending. Put the scheduler's reason in /root/csch/pressure.txt.
  3. Put the taint lab=only:NoSchedule on the node, compare no-tol, which has no toleration, with with-tol, which has one, and put it in /root/csch/taint.txt.
  4. Label the node disktype=ssd and create three Pods, aff-required (one that matches), aff-impossible (a label that does not exist, as required), and aff-preferred (a label that does not exist, as preferred), and put it in /root/csch/affinity.txt.
  5. Put a required anti-affinity on the spread Deployment (2 replicas), and put in /root/csch/anti.txt that only one comes up because there is only one node.
  6. Create the PriorityClasses low-prio and high-prio, fill the node with the low one, then add important (the high one) and put the preemption that occurs in /root/csch/preempt.txt.
  7. On the even Deployment, set topologySpreadConstraints to ScheduleAnyway and put the result in /root/csch/spread.txt.
  8. In /root/csch/report.md, write three lines, pending_reason=Insufficient, taint_effect=NoSchedule, and high_priority_value=, and an explanation.

Reference

What the node has, and how much

Put the node's capacity and allocatable, and how much is requested right now, in /root/csch/capacity.txt.

capacity is what the hardware has, allocatable is what is left after the system's share, and Allocated resources is the sum already requested.

When there is not enough room, it waits

Make three Pods, hog-a, hog-b, hog-c, compete for room. The first two must be placed and the third must be Pending. Put the scheduler's reason in /root/csch/pressure.txt.

Set a large CPU requests so that only two fit. The only thing the scheduler looks at is requests.

The node refuses

Put the taint lab=only:NoSchedule on the node, compare no-tol, which has no toleration, with with-tol, which has one, and put it in /root/csch/taint.txt.

A taint is what a node sets, and a toleration is what a Pod has. NoSchedule blocks only new placements.

required and preferred

Label the node disktype=ssd and create three Pods, aff-required (one that matches), aff-impossible (a label that does not exist, as required), and aff-preferred (a label that does not exist, as preferred), and put it in /root/csch/affinity.txt.

With required, if it cannot be matched it stays Pending forever, and with preferred, it is placed even if it cannot be matched.

Avoid the same node

Put a required anti-affinity on the spread Deployment (2 replicas), and put in /root/csch/anti.txt that only one comes up because there is only one node.

topologyKey: kubernetes.io/hostname means 'they just need to be on different nodes.' With one node, there is nowhere to place the second.

A higher priority takes the room

Create the PriorityClasses low-prio and high-prio, fill the node with the low one, then add important (the high one) and put the preemption that occurs in /root/csch/preempt.txt.

If you fill the node with low priority and then add a high-priority Pod, the scheduler evicts the low one.

Spread evenly

On the even Deployment, set topologySpreadConstraints to ScheduleAnyway and put the result in /root/csch/spread.txt.

There is only one node, so it has to be whenUnsatisfiable: ScheduleAnyway. With DoNotSchedule, nothing comes up.

What you learned

In /root/csch/report.md, write three lines, pending_reason=Insufficient, taint_effect=NoSchedule, and high_priority_value=, and an explanation.

Write the order for narrowing down Pending together with the three lines pending_reason=, taint_effect=, and high_priority_value=.