CNPE — Cloud Native Platform Engineer
Read quotas, defaults and QoS as different measurements
One-line summary
ResourceQuota is the upper bound a tenant can use, requests are an input to scheduling, and actual usage is an observed value. Do not treat the cap, the requests, and real usage as the same number.
Why this was needed
Suppose a practice cluster has 24 cores of allocatable CPU and three tenants have requests.cpu quotas of 12, 9, and 9 cores. The sum of the caps is 30 cores, and the ratio to capacity is 1.25. This does not mean 30 cores are reserved right now, nor that CPU is being used at 125%. It is a planning signal that the cluster may not be able to accommodate every team requesting its full cap at the same time. You must check separately the sum of requests of the Pods actually placed and the usage.
Adding up the per-team caps is only a starting point. A workload concentrated on a particular node pool or storage zone may fail to start even when there is spare capacity overall, and you also need headroom for a single node failure. Do not decide how many servers to buy by adding up the quota hard values alone.
How it works
Admission review and seat assignment
ResourceQuota checks namespace-scoped resource and object-count limits at admission. It is not a mechanism that continuously measures real CPU usage. You can also limit the number of CRs with a key such as count/appclaims.platform.labhub.io. RBAC, policies, and separate platform logic can also restrict issuance, so quota is not the only means. Check what is counted in the official ResourceQuota documentation.
LimitRange handles per-object minimums, maximums, defaults, and so on. default is the default for limit, and defaultRequest is the default for request. It applies when a Pod is created, and even if you change the policy, existing Pods are not fixed automatically. Do not give conflicting defaults through several LimitRanges. Based on the official LimitRange documentation, get into the habit of comparing the file before submission with the object after admission.
Guaranteed has conditions
This explanation uses an example that sets resources per container and does not declare Pod-level resources. To be Guaranteed, every regular and init container must have a request and a limit for each of CPU and memory, and they must be equal. Even if only memory is filled with equal values, a missing CPU condition means the Pod is not Guaranteed. If there is at least one resource declaration but the Guaranteed conditions are not met, the Pod is Burstable. Compare the official QoS lab with the final status.qosClass.
| Final resources of the example container | QoS | Reason |
|---|---|---|
| CPU 100m/100m, memory 64Mi/64Mi | Guaranteed | Request and limit are equal for both resources |
| No CPU, memory 64Mi/64Mi | Burstable | The CPU condition is missing |
| CPU 25m/100m, memory 32Mi/64Mi | Burstable | The request is smaller than the limit |
| No CPU or memory declared on any container | BestEffort | Neither resource has a request or limit |
If a container specifies a limit and omits the corresponding request, the request is set to the limit value. Distinguish this case from the case where both are omitted and the LimitRange defaults apply. A request that is already written is not something the default overwrites. When you read the official memory defaults example, try predicting the outcome for each of these two inputs by swapping them.
Reading exercise: look at the stored object, not the input
The example below is meant to be run only on your own dedicated learning cluster, not on a billed or production cluster. If the cnpe-default-demo namespace already exists, do not reuse it; choose a new name. First predict the Pod that receives the two defaults, and then read the actual result.
kubectl create namespace cnpe-default-demo
kubectl -n cnpe-default-demo apply -f - <<'YAML'
apiVersion: v1
kind: LimitRange
metadata:
name: defaults
spec:
limits:
- type: Container
default:
cpu: 100m
memory: 64Mi
YAML
kubectl -n cnpe-default-demo run sample --image=busybox:1.36 -- sleep 3600
kubectl -n cnpe-default-demo get pod sample -o jsonpath='{.spec.containers[0].resources}'
kubectl -n cnpe-default-demo get pod sample -o jsonpath='{.status.qosClass}'
The expectation is a request and limit of CPU 100m and memory 64Mi, and Guaranteed. An image pull failure and a failure to apply resource defaults are separate things. To observe the container actually running, you need an environment with image access and a real kubelet. Do not use kwok's Running marker as evidence of actual execution.
The counterexample is to create a new Pod in a separate namespace with only the CPU default removed. Check that it is Burstable even though the memory request and limit are equal. Changing only the policy in the original namespace and then reading an existing Pod is not this counterexample. After you record the results, delete the namespace that holds only the Pods and policies you created.
kubectl delete namespace cnpe-default-demo
What it looks like in the field
Suppose that in a hypothetical cost meeting someone looks only at a workload with a 4-core request and 0.3-core real usage and immediately lowers the request to 0.3. If peak hours, startup cost, and failure headroom were not examined, an outage may arrive before any savings. Look at the usage distribution over time periods and the SLO, make the change in a small scope, and then check scheduling, latency, and restarts together. The QoS class itself does not guarantee performance or zero downtime.
If quota usage is empty or a status unknown error appears, check the object's status, whether the CRD is Established, API discovery, and the controller state. Deleting and recreating unconditionally can briefly remove admission protection. Act after you confirm the cause and the scope of recovery, and test that new requests are allowed and rejected respectively.
Observed case: you got in, but there is no seat
On 2026-09-11, we actually confirmed the following on a separate single-node k3s v1.36.4+k3s1 VM. It was not a load test of a production service but an experiment that separated admission from scheduling using a small sleep container. We did not generate load to use 9 cores of CPU; we requested 9 cores.
| Field checked | Observed value | Meaning |
|---|---|---|
| Node status.allocatable.cpu | 8 | CPU capacity for Pods on this node |
| ResourceQuota status.hard.requests.cpu | 10 | Request cap for this Namespace |
| Pod spec.containers[0].resources.requests.cpu | 9 | CPU requested by the single container |
| ResourceQuota status.used.requests.cpu | 9 | The request of a Pod that is not yet running is also counted |
| Pod status.phase / spec.nodeName | Pending / none | Stored in the API but not placed on a node |
| PodScheduled condition | False, Unschedulable, Insufficient cpu | The scheduler reports a CPU shortage |
A request of 9 fits within the quota of 10, so the object was created, but a request of 9 could not be placed on a node with 8. Raising the quota to 20 at this point does not make the node bigger. If every node has 8 cores, simply adding nodes of the same size will not let this single Pod in either. You must verify whether the request was sized wrongly, or consider design changes such as a sufficiently large node or splitting the application. Compare the case of a Pod larger than the node in the official resource troubleshooting documentation.
In the experiment, we deleted the Pod with the same name and recreated it with a request of 25m. A new UID was issued, Running, Ready, and an actual exec were confirmed, and the quota cap stayed as it was. This is evidence of restored ability to run, not evidence of a right-sized request or of cost savings. Do not apply the 25m of a sleep container directly to a real order server. CPU usage time series, throughput, latency, and error rate were not measured in this experiment.
Conversely, in another Namespace, after filling a 100m quota with two Pods requesting 50m each, creating a third 50m Pod made the API return Forbidden and exceeded quota, and the third object was not stored. In this case there is no Pod for the scheduler to look at. If a Deployment fails to create a Pod, you must also check the events of the parent object. If you look only for Pending Pods, you miss admission failures.
Test the tenant boundary for reads and changes separately
In another experiment on the same VM, we granted a ServiceAccount only pods get/list in its own Namespace. When we sent real queries as that account, it could read its own Pods, but querying Pods in another Namespace and modifying its own ResourceQuota were Forbidden. The quota the administrator read again after the rejection was still 100m. This is why you should look together at the allowed request, the rejected request, and the unchanged state, rather than merely noting that a Role YAML exists.
This experiment is an API permission check performed by the administrator impersonating with --as. It does not verify actual login, token issuance, network blocking, or kernel isolation. If you configure only a quota and give the tenant permission to edit the quota, the tenant can raise the cap themselves. Conversely, blocking API reads does not mean that traffic to another team's service port is impossible. Read the Kubernetes multi-tenancy documentation while distinguishing the control plane from the data plane.
Judge for yourself. ① Quota used is 9, so are 9 cores of CPU running? ② Is the Pod restored under the same name the same object as before? ③ If another team's read is rejected, is network isolation also done? The answer to all of these is no, and the grounds are, respectively, how requests are counted, the change of UID, and the difference in the scope of the test.
What to check in the next quiz
Judge what 24 and 30 are sums of, the QoS of a Pod that limits only memory, and the effect of a policy change on existing Pods. The goal of this unit is not to stretch the success of a single healthy example into evidence of overall tenant isolation or cost optimization.