CNPE — Cloud Native Platform Engineer
Multi-Tenant Resources and Incident Response
Goal
You create three tenants and divide resources with quotas and LimitRanges, compute and record the overcommit ratio yourself, reproduce a scheduling incident and classify it with numbers, recover without deleting the requirement, and create the platform's own recording rule and alert and upload them to the cluster.
Why it matters
What you practice here is not only tool operation but judgment based on evidence. That is the order: what you count first when an incident happens, where you write down that value, and what you will call recovery.
The last item is especially important. Making a Pod start by deleting a constraint and making it start by satisfying the constraint are indistinguishable on the screen. Both are Running and both silence the alert. It is a person who distinguishes them, and getting that judgment into your hands is the purpose of this lab.
This cluster has three nodes of 8 cores each. So the cluster's total allocatable CPU is 24 cores. The overcommit ratio is the quota total divided by this value.
This environment uses a real API server, a real scheduler, and KWOK virtual nodes. Running is for learning scheduling state and is not evidence that a GPU or a ledger process actually ran. Prometheus Operator and Alertmanager are not run. The PromQL in the last two steps is actually evaluated with promtool 3.0.1 included in the image, and the CR is verified down to what the API stores.
The working directory is /root/cnpe-ops. The time-series tests use virtual time, so you do not actually wait 25 minutes. Each grading runs a check within 45 seconds and does not modify student files or API objects.
Steps
- Create the namespaces
tenant-red,tenant-green, andtenant-goldand attach theplatform.labhub.io/tenantlabel to each. The value is the name withtenant-removed. Then write two lines,tenantsandtenant_list, in/root/cnpe-ops/inventory.txt. Join the list in name order with commas only. - In
/root/cnpe-ops/quotas.yaml, write the ResourceQuota for the three tenants. The names are each the namespace name with-quotaadded, liketenant-red-quota, and you must includerequests.cpu. Make the sum of the three quotas'requests.cpudivided by 24 greater than 1.00 and at most 1.25, and do not bring any tenant below 6 cores. Then write three lines,quota_cpu,allocatable_cpu, andovercommit, in/root/cnpe-ops/capacity.txt. - In
/root/cnpe-ops/limitranges.yaml, write the LimitRange for the three tenants. Name them liketenant-red-limits. Include bothdefaultRequestanddefault, but the CPU ofdefaultRequestmust be lower than the CPU ofdefault, and usemaxto prevent a single container from requiring 4 CPU cores. - In
/root/cnpe-ops/ledger.yaml, write the Deploymentledgerand apply it totenant-red. replicas is 3, the Pod label isapp=ledger, thenodeSelectoris the singleplatform.labhub.io/pool=gpu, the container requests are CPU 500m and memory 512Mi, and the image is pinned by digest. At this point the Pods cannot start. - In
/root/cnpe-ops/triage.txt, write four lines,selector_key,selector_value,replicas, andrequired_cpu. Forrequired_cpu, write the container request multiplied by the replica count, in millicores. - Recover by attaching the
platform.labhub.io/pool=gpulabel to only one node. You must leave the Deployment'snodeSelectorand replicas as they are. Check until all three Pods are Running. - In
/root/cnpe-ops/platform-rules.yaml, put the groupplatform-sloandinterval: 1m. The recording ruleplatform:provision_success:ratio5mis the sum of the success request rates over the last 5 minutes / the sum of the total request rates. Compute the rate per series and then sum, and if there is only a failure series, it is 0%, if there is only a success series, it is 100%, and for no traffic or no observation, do not emit a ratio series.PlatformProvisionFailingfires when this ratio stays below 95% for 10 minutes and resolves when it recovers. Setseverity: critical, a meaningfulsummary, and arunbook_url. After the syntax check, pass the 13 time-series events withpython3 /opt/fixtures/cnpe_slo_contract.py rules. 95% is this lab's threshold and not a recommended production SLO. - Create the namespace
monitoring, and in/root/cnpe-ops/platform-rules-cr.yaml, write the PrometheusRuleplatform-sloand apply it. The rules file, the manifest's spec.groups, and the API's spec.groups must match in their entirety. In/root/cnpe-ops/ops-report.txt, record four lines without duplicates,tenants,pool_nodes,ledger_running, andalert_rules, as integers you queried. The targets in this lab are 3, 1, 3, and the actual number of alert rules, respectively. Check withpython3 /opt/fixtures/cnpe_slo_contract.py deploy. A successful API store does not mean that the production Prometheus loaded the rules or that alert delivery succeeded.
Reference
- The value you count by label is the value the platform itself knows. Count with
kubectl get ns -l platform.labhub.io/tenant. - You can recount a node's allocatable with
kubectl get nodes -o json | jq '[.items[].status.allocatable.cpu | tonumber] | add'. - To see what a LimitRange actually fills in, the surest way is to submit a Pod with its requests left empty using
--dry-run=serverand read the result. - After attaching the label, give the scheduler time to go around once. If you check and judge immediately, you end up reverting a correct action.
- One common mistake is deleting the
nodeSelectorto make the Pod start. That is not a recovery but removing a requirement. - Another is uploading a PrometheusRule whose content differs from the rules file you verified. Then passing promtool guarantees nothing.
- Read Prometheus rule unit testing and alerting rules. Syntax and firing behavior are separate.
- runbook.example.invalid is an address for illustrating the format. In production, replace it with a document that has the real owner, diagnosis, and recovery procedure.
Make tenants countable by machine
Create the namespaces tenant-red, tenant-green, and tenant-gold and attach the platform.labhub.io/tenant label to each. The value is the name with tenant- removed. Then write two lines, tenants and tenant_list, in /root/cnpe-ops/inventory.txt. Join the list in name order with commas only.
If there is only a naming convention and no label, the tenant list exists only in people's memory. After attaching the labels, recount with a label selector and write down that result.
Separate the promised amount from the amount actually held
In /root/cnpe-ops/quotas.yaml, write the ResourceQuota for the three tenants. The names are each the namespace name with -quota added, like tenant-red-quota, and you must include requests.cpu. Make the sum of the three quotas' requests.cpu divided by 24 greater than 1.00 and at most 1.25, and do not bring any tenant below 6 cores. Then write three lines, quota_cpu, allocatable_cpu, and overcommit, in /root/cnpe-ops/capacity.txt.
A quota is applied independently per namespace and does not reference cluster capacity. So a person must count the total separately. Three nodes of 8 cores each make the denominator 24.
Decide the values to fill in for those who did not write them
In /root/cnpe-ops/limitranges.yaml, write the LimitRange for the three tenants. Name them like tenant-red-limits. Include both defaultRequest and default, but the CPU of defaultRequest must be lower than the CPU of default, and use max to prevent a single container from requiring 4 CPU cores.
default is the default for limit, and defaultRequest is the default for request. A request that is already declared is not overwritten by the default. In this task you set the default CPU request smaller than the limit to tell the reserved amount apart. Whether it is Guaranteed requires checking the CPU and memory conditions of all containers together. max is the per-container upper bound.
Reproduce the incident
In /root/cnpe-ops/ledger.yaml, write the Deployment ledger and apply it to tenant-red. replicas is 3, the Pod label is app=ledger, the nodeSelector is the single platform.labhub.io/pool=gpu, the container requests are CPU 500m and memory 512Mi, and the image is pinned by digest. At this point the Pods cannot start.
Right now these nodes have no pool label. So this workload cannot start. That state is the normal state for this step. What is graded is not the Pod's current state but what the workload requires.
Write the result of the classification as numbers
In /root/cnpe-ops/triage.txt, write four lines, selector_key, selector_value, replicas, and required_cpu. For required_cpu, write the container request multiplied by the replica count, in millicores.
Write four numbers, not impressions. The total CPU required is the container request multiplied by the replica count, written in millicores. Read the values directly from the cluster.
Recover without deleting the constraint
Recover by attaching the platform.labhub.io/pool=gpu label to only one node. You must leave the Deployment's nodeSelector and replicas as they are. Check until all three Pods are Running.
If you delete the selector, it starts immediately, but that is not a recovery but removing a requirement. Fix the node side to satisfy the constraint. And add only as many as needed. If you attach it to two, the meaning of that pool gets blurred.
Write the platform's own recording rule and alert
In /root/cnpe-ops/platform-rules.yaml, put the group platform-slo and interval: 1m. The recording rule platform:provision_success:ratio5m is the sum of the success request rates over the last 5 minutes / the sum of the total request rates. Compute the rate per series and then sum, and if there is only a failure series, it is 0%, if there is only a success series, it is 100%, and for no traffic or no observation, do not emit a ratio series. PlatformProvisionFailing fires when this ratio stays below 95% for 10 minutes and resolves when it recovers. Set severity: critical, a meaningful summary, and a runbook_url. After the syntax check, pass the 13 time-series events with python3 /opt/fixtures/cnpe_slo_contract.py rules. 95% is this lab's threshold and not a recommended production SLO.
A vector with value 0 also activates an alert. If you use a bool comparison directly in an alert, a series remains even in a healthy state. Distinguish the empty numerator when there are only failures from the denominator when there is no total traffic. First read the hypothetical event and time point in the failure message, and then check the formula and for.
Upload the verified rules and record the current state
Create the namespace monitoring, and in /root/cnpe-ops/platform-rules-cr.yaml, write the PrometheusRule platform-slo and apply it. The rules file, the manifest's spec.groups, and the API's spec.groups must match in their entirety. In /root/cnpe-ops/ops-report.txt, record four lines without duplicates, tenants, pool_nodes, ledger_running, and alert_rules, as integers you queried. The targets in this lab are 3, 1, 3, and the actual number of alert rules, respectively. Check with python3 /opt/fixtures/cnpe_slo_contract.py deploy. A successful API store does not mean that the production Prometheus loaded the rules or that alert delivery succeeded.
If only the group name is the same but the expression, threshold, or for differs, it is not the same rule. Wrap the verified file as-is into the CR and then compare the whole spec.groups in the API. Do not record a failed query as 0; resolve the API error first.