KCNA — Kubernetes and Cloud Native Associate
Low usage is not the same as available scheduling capacity
One-line summary
requests are a promise used to choose a place, and limits are a setting that caps how much can be used while running. You do not conclude that there is room for a new Pod just because CPU utilization is low.
Why this was needed
Only five people are sitting in a classroom, yet a new student cannot get in. The seats look empty, but the remaining chairs are reserved for the afternoon class. The number of people sitting now and the number of reserved seats are different ledgers. In Kubernetes too, current usage and the resource requests used for scheduling are different. The reservation metaphor is for understanding placement decisions; it does not mean that one CPU core is physically dedicated to a particular container.
Suppose your team deployed a small Python app and the Pod stays in Pending. The code is merely asleep, so it uses almost no CPU. But the manifest mistakenly says requests.cpu is 8, and every node is smaller than that. How idle the app is does not solve this problem. This is because the request declared before the app starts has already exceeded the size of any node it could be placed on.
How it works
First, distinguish four numbers.
| Number | The question it answers | Where to check |
|---|---|---|
| capacity | How much total resource does the node report? | Node.status.capacity |
| allocatable | How much resource is available for placing Pods? | Node.status.allocatable |
| requests | How much request is counted when placing this Pod? | The container's resources.requests |
| Actual usage | How much was used in a particular observation window? | The resource metrics API and monitoring |
This lesson focuses on the CPU and memory requests of ordinary containers. If init containers, Pod overhead, Pod-level resource settings, or extended resources are involved, you must check the calculation rules further. Do not apply the arithmetic of a simple single container as is to every workload.
For example, if a node's allocatable CPU is 2 and the sum of requests already placed is 1.4, adding a new request of 800m makes a total of 2.2. It does not fit on the CPU condition alone. Even if actual usage is observed as 200m, the sum of requests does not shrink by itself. Conversely, the sum of requests fitting does not guarantee successful placement. Other conditions such as memory, node selection rules, taints, and volumes must also be satisfied.
CPU 1 is one CPU unit and equals 1000m. 250m is 0.25 CPU, not 250 percent of the whole node. The Mi and M of memory are not the same either. 64Mi is 64×1024×1024 bytes, and 64M is 64×1000×1000 bytes. If you omit units when reading a resource problem, you end up comparing different ledgers with the same number.
limits are a different axis. A CPU limit appears as throttling that restricts execution time, and a memory limit can appear as an OOM termination depending on the situation. Exceeding requests does not cause an immediate termination. If you write a low request and a high limit, the container may use more when there is spare room. That does not mean lowering requests unconditionally is the solution. Under-requesting can send too many workloads to the same node and create performance contention and pressure.
What it looks like in practice
You practice the order for narrowing down the failure point.
kubectl -n kcna-capacity get pod oversized -o json
kubectl -n kcna-capacity describe pod oversized
kubectl get nodes -o json
First, check that the Pod object actually exists. Second, look at whether spec.nodeName is set and at the PodScheduled condition. Third, read the FailedScheduling event that corresponds to that Pod UID. If you pull in an event from a past Pod with the same name, you may mistake an error from before the replacement for the current cause. Pending also includes other preparation steps such as image download, so you do not conclude it is a CPU shortage from that one word alone.
Once Insufficient cpu is confirmed, compare the declared request with the node's allocatable and the existing requests. If the request is a size that is truly needed, consider options such as a larger node, splitting the work, or cleaning up unnecessary reservations. If it is a typo, fix it with evidence. Deleting another team's running Pods or editing node state just to make the numbers fit is not diagnosis.
What you will do in the lab that follows
You declare a CPU request larger than the real allocatable of your personal k3s, but you do not run a load program that burns CPU. After collecting evidence that the Pod stored in the API is not placed, you reclaim only the test Pod with that UID and compare whether a small Pod with a 50m request starts. You preserve the healthy comparison group in a separate namespace to the end. This experiment is different from one that directly measures CPU saturation, throttling, or OOM.