KCNA — Kubernetes and Cloud Native Associate
The accounting behind defaults, quotas, and admission
One-line summary
A LimitRange handles the defaults and ranges of individual objects, and a ResourceQuota handles the total of a namespace. A request that runs into a quota can be rejected at creation time rather than the Pod waiting in Pending.
Why this was needed
If one small team accidentally creates thousands of objects, other teams in the cluster are affected too. Resource policies limit not only malicious attacks but also the damage a typo in a single loop can cause. However, the mere fact that you put a limit in place does not guarantee the behavior you want. You have to actually check which requests it applies to, whether it touches existing objects, and what the usage counts.
Suppose you tried to create a new Pod and got Forbidden. If you always treat this as an account permission problem, you end up with the wrong response of granting more RBAC permissions. Even with the same HTTP status, the cause in the response body and the related object matter. A quota overrun is not solved by adding permissions. Making it succeed by deleting the limit is not a solution that understands the policy either.
How it works
In this experiment we set, for a single container, a default request of CPU 100m and memory 32Mi and a default limit of CPU 200m and memory 64Mi. You read side by side a new Pod created with resources omitted from the input and the Pod stored in the API. If values appear on the stored object, they may be defaults applied during the admission process, not values from the user's file. Do not assume that an existing Pod that was created with CPU 50m explicitly was changed to 100m.
The requests.cpu of a ResourceQuota limits the sum of the CPU requests of the target Pods that belong to this space. It is not a number measuring what percentage of CPU is actually being used right now. In this example the ledger moves as follows.
| State | Sum of stored requests | Allowed cap | Result for a new request |
|---|---|---|---|
| fit 50m + defaulted 100m | 150m | 200m | Adding extra 100m makes 250m, so it is rejected |
| After raising the cap | 150m | 300m | The same extra request comes in and the total is 250m |
| Lowering the cap again | 250m | 100m | Existing Pods are kept, and a new small request is rejected |
| Reclaiming extra and readjusting the cap | 150m | 200m | small 50m comes in and the total is 200m |
The cap changes in this table are an experiment for learning the meaning of policy in a disposable namespace. In production, an authorized person who has reviewed usage, business needs, and team allocation changes the policy. This is not guidance to raise a quota simply to make an error go away.
Lowering a quota below current usage does not automatically terminate existing Pods. So you can observe a state where status.used is greater than spec.hard. This is not evidence that the policy is invalid. You have to check whether it keeps the existing objects while rejecting additional creation. If you need to force usage down, you have to make a separate operational decision about which work to terminate.
A ResourceQuota may not limit just one axis. It can check requests.cpu, requests.memory, limits.cpu, limits.memory, and the number of Pods together, so you do not assert success by counting only CPU. This experiment leaves headroom in the other caps to isolate the effect of the CPU request. You also distinguish whether an individual cap of the LimitRange was violated or the total cap of the ResourceQuota.
What it looks like in practice
You look at the input before creating the Pod, the API response, and the lookup result after creation as one bundle of evidence. For a quota rejection, check the actual HTTP 403, the Forbidden reason in the Kubernetes Status, and a description that includes team-budget and requests.cpu. Then check that no Pod of the same name exists. Do not record that the desired policy worked when what you received was an account permission error or a TLS error.
Conversely, HTTP 201 means the object creation was accepted, not that the app has finished starting. Continue to check that the new Pod UID matches the response and that the Pod is placed on a node and is Running and Ready. Scheduling and admission are different stages, so "this request is within the quota" and "there is a node where this Pod can run" are separate.
kubectl -n kcna-capacity get limitrange defaults -o json
kubectl -n kcna-capacity get resourcequota team-budget -o json
kubectl -n kcna-capacity get pods -o json
The status ledger may still show the old value right after a change. After deleting the extra Pod you own, confirm the actual absence of the object and the decrease in quota usage, and then send the next request. Do not assume completion from a fixed short sleep alone. Put a time limit on the waiting, and if the time is exceeded, preserve the observation rather than papering it over as a success.
What you will do in the next lab
You directly compare two failures that occur even when resources seem to be left over. You prove the placement failure with the actual Pod UID and scheduler events, and the admission rejection with the API state and the absence of the object. After the quota is reduced, you check that the UIDs and containers of the existing fit, defaulted, and extra Pods are unchanged, then reclaim only the extra you investigated and have a small request newly accepted. You do not change the healthy comparison group in the other namespace.