KCNA — Kubernetes and Cloud Native Associate
CPU looks idle, so why can the Pod not get in?
Goal
You distinguish, on a real k3s, a scheduling failure caused by resource requests from an API creation rejection caused by a quota.
Why it matters
Even if current utilization is low, the declared request may not fit on the node. Even if admission is allowed, actual startup is a separate stage. Without deleting a policy or reclaiming someone else's Pod to get through, you compare defaults, totals, identities, and actual responses. The experiment is a small idle app on a personal VM, not a CPU saturation, throttling, or OOM load experiment. Do not bring in a production kubeconfig or real secrets. This is a 50-minute lab. If needed, extend before it expires, and remember that the VM and files are reclaimed when the session ends.
Prepared environment and helpers
kcna-capacity is the experiment space, and kcna-capacity-control/control is the healthy comparison group. Both spaces have the restricted/v1.36 policy. The app runs with a pinned digest, Always pull, non-root, cap drop ALL, RuntimeDefault, a read-only root, and the ServiceAccount token not mounted. Do not change the node or the comparison group. Keep all of the student's files under /root/kcna-resources. The student writes the input files, and the observation files are made by capture, which reads the actual API and Pod state. Do not fabricate success, UIDs, or events. The helper is python3 /opt/fixtures/kcna_resources_lab.py followed by act, capture, grade, or prepare and a step number. observe only reads the current state. grade also changes neither student files nor Kubernetes objects. act's API response is preserved in /opt/fixtures/kcna-resource-action-N.json. Do not repeat the same request to recreate a past 403 or 201.
Steps
- Save /root/kcna-resources/baseline.json with capture 1. Investigate the Node's actual allocatable CPU and UID, the empty kcna-capacity namespace, and the UID, container ID, Ready, and restart count of kcna-capacity-control/control.
- Write requests and limits objects in oversized-resources.json. For both cpu values, write a value 1 CPU larger than the actual node allocatable as a string in m units, and memory is 32Mi and 64Mi respectively. Save unschedulable.json with act 2 and capture 2. The oversized Pod must exist but be unplaced and Pending, and there must be an Unschedulable condition and an Insufficient cpu event for that UID.
- Write requests={cpu:50m,memory:32Mi}, limits={cpu:200m,memory:64Mi} as JSON in fit-resources.json. act 3 reclaims only the oversized UID you investigated and creates a small fit Pod. Save fitted.json with capture 3 and check a different Pod UID and normal startup.
- Write type=Container, defaultRequest={cpu:100m,memory:32Mi}, default={cpu:200m,memory:64Mi} in defaults-policy.json. Save defaults.json with act 4 and capture 4. Compare the actual 100m request of the defaulted Pod, whose resources were omitted, with the preserved 50m of the existing fit.
- Write requests.cpu=200m, requests.memory=128Mi, limits.cpu=1, limits.memory=256Mi, pods=4, all as strings, in quota.json. act 5 creates team-budget and then requests the creation of an extra with resources omitted. Save quota_denied.json with capture 5. With usage at 150m, the additional 100m must be rejected with an actual quota HTTP 403, and extra must not exist.
- Write requests.cpu=300m in increase.json. Save quota_admitted.json with act 6 and capture 6. Check the actual HTTP 201 of the same extra creation request, the new Pod UID and Ready, and usage of 250m. Do not change the other quota axes or the existing Pods.
- Write requests.cpu=100m in lower.json. Save quota_lowered.json with act 7 and capture 7. The existing Pods with usage of 250m must be kept with the same UID and containers, and a new 10m small request must be rejected with an actual HTTP 403. Do not interpret the policy reduction as an automatic termination of existing Pods.
- Write remove=extra, requests.cpu=200m, replacement={requests:{cpu:50m,memory:32Mi},limits:{cpu:200m,memory:64Mi}} in recover.json. act 8 reclaims only the extra you investigated, and after usage decreases, adjusts the budget and creates small. Save budget_restored.json with capture 8. Prove the actual HTTP 201, Ready, the final request total of 200m, and the preservation of the comparison group.
Notes
Notation such as cpu:50m in the text explains fields and values. The file you save must be valid JSON in which keys and strings are wrapped in quotes, as in the examples. CPU 1 is 1000m. Step 2 first reads this VM's actual allocatable to compute the value, so do not memorize the value of a particular VM. Completed inputs and observations are preserved. If you wrote an input wrong while in progress, read the cause of the error and fix it, but do not rewrite past observations you have saved. Preparation fills only the earlier steps that are missing. It does not create the current task's inputs and observations or overwrite existing inputs you wrote incorrectly. Grading has a 60-second limit and step preparation a 90-second limit. Budgets and usage are checked for convergence, and success is not assumed from a fixed sleep alone. If an object was created right after an interruption but there is no identity record, it fails rather than taking over an arbitrary object. It does not erase the learning record with an automatic reset. The record hash is for preventing mistakes and is not a remote attestation device that stops malicious changes by root. Official sources: Resource management · Defaults policy · Quotas.
Investigate the placement budget and the healthy comparison group
Save /root/kcna-resources/baseline.json with capture 1. Investigate the Node's actual allocatable CPU and UID, the empty kcna-capacity namespace, and the UID, container ID, Ready, and restart count of kcna-capacity-control/control.
A usage graph and allocatable answer different questions.
A request that cannot get in even though it looks idle
Write requests and limits objects in oversized-resources.json. For both cpu values, write a value 1 CPU larger than the actual node allocatable as a string in m units, and memory is 32Mi and 64Mi respectively. Save unschedulable.json with act 2 and capture 2. The oversized Pod must exist but be unplaced and Pending, and there must be an Unschedulable condition and an Insufficient cpu event for that UID.
Do not look only at Pending; look at nodeName, the PodScheduled condition, and events of the same UID together.
Start with a request of a fitting size
Write requests={cpu:50m,memory:32Mi}, limits={cpu:200m,memory:64Mi} as JSON in fit-resources.json. act 3 reclaims only the oversized UID you investigated and creates a small fit Pod. Save fitted.json with capture 3 and check a different Pod UID and normal startup.
Do not change node resources or delete the comparison group.
Apply defaults to omitted resources
Write type=Container, defaultRequest={cpu:100m,memory:32Mi}, default={cpu:200m,memory:64Mi} in defaults-policy.json. Save defaults.json with act 4 and capture 4. Compare the actual 100m request of the defaulted Pod, whose resources were omitted, with the preserved 50m of the existing fit.
Look at whether a value that was not in the input appeared on the stored object, and compare the existing object too.
A quota that rejects Pod creation itself
Write requests.cpu=200m, requests.memory=128Mi, limits.cpu=1, limits.memory=256Mi, pods=4, all as strings, in quota.json. act 5 creates team-budget and then requests the creation of an extra with resources omitted. Save quota_denied.json with capture 5. With usage at 150m, the additional 100m must be rejected with an actual quota HTTP 403, and extra must not exist.
Tell from the response body whether the cause of Forbidden is RBAC or team-budget.
Distinguish admission from actual execution
Write requests.cpu=300m in increase.json. Save quota_admitted.json with act 6 and capture 6. Check the actual HTTP 201 of the same extra creation request, the new Pod UID and Ready, and usage of 250m. Do not change the other quota axes or the existing Pods.
201 is acceptance of the creation. Check that the same UID is actually running.
Even if you lower the cap, existing Pods remain
Write requests.cpu=100m in lower.json. Save quota_lowered.json with act 7 and capture 7. The existing Pods with usage of 250m must be kept with the same UID and containers, and a new 10m small request must be rejected with an actual HTTP 403. Do not interpret the policy reduction as an automatic termination of existing Pods.
Even if the quota's used is greater than hard, existing objects can be kept.
Reclaim only the investigated request and restore the budget
Write remove=extra, requests.cpu=200m, replacement={requests:{cpu:50m,memory:32Mi},limits:{cpu:200m,memory:64Mi}} in recover.json. act 8 reclaims only the extra you investigated, and after usage decreases, adjusts the budget and creates small. Save budget_restored.json with capture 8. Prove the actual HTTP 201, Ready, the final request total of 200m, and the preservation of the comparison group.
Check not only the deleted name but also the convergence of the UID and the usage ledger.