CNPE — Cloud Native Platform Engineer
Fixing misleading metrics in a provisioning API
Goal
You declare the SLI's denominator, fix the instrumentation code that drops failures, and compare real HTTP responses with the metrics. After telling duplicate collection apart from retries, you use the error budget of three independent observation windows to decide on a general feature deployment.
Why it matters
Instrumentation that counts only the success path can show an ever better success rate as failures increase. Even a correct PromQL cannot fix raw data that was collected wrongly. Before you write alert rules, this lab checks what is counted as one event and whether any evidence is missing. You do not pass merely by matching function names; you are checked with varied inputs, real HTTP execution, and counterexamples you design yourself.
This lab actually runs a Python server inside a Pod. It does not use the Kubernetes API or KWOK's Running state as execution evidence. It does not deal with external systems, a real GPU, production Prometheus scraping, or real deployment. Latency is server processing time and differs from the client's network round-trip time. The numbers of the 30-day window are hypothetical data for learning.
The working path is /root/cnpe-sli. Python uses only the standard library. When the session ends, the files and diagnostic records disappear, so keep the deliverables you need somewhere else. The helper is /opt/fixtures/cnpe_sli_workbench.py, and each grading reads the files and terminates the temporary process.
Steps
- Declare the unit, the denominator, and the targets in scope.json.
- Classify the status code and path with classify in classifier.py.
- Complete the classification of slow success responses in the same function.
- Replay real HTTP and save http-receipt.json.
- With summarize.py, remove collection duplicates but count real retries separately.
- With budget.py, handle observation gaps, fractional budgets, and the exhaustion boundary.
- With report.py and report.json, combine the availability and latency judgments of the three windows.
- With counterexamples.json, distinguish five wrong implementations.
Reference
- The step card has the exact function signature and the JSON field contract. The file examples show only the format and are not answers.
- You can check each step directly with
python3 /opt/fixtures/cnpe_sli_workbench.py grade /root/cnpe-sli 단계번호(the placeholder stands for the step number). exercise /root/cnpe-sliuses a temporary port to actually produce healthy, 1.1-second delay, 503, 504, 400, and 404. X-Lab-Scenario is a failure-selection header for learning and is not a production feature.- Only in this task are 2xx and 5xx defined as valid attempts. If you unconditionally exclude 429 or 4xx in production, you can hide failures that are the service's responsibility. You must agree on the contract for what a valid request is.
- Read SLI design, instrumentation principles, and an error budget policy example together.
Leave the denominator as a contract first
Create /root/cnpe-sli/scope.json. route is /provision, unit is attempt, eligible_classes is [2,5], latency_limit_ms is 1000, availability_target is 0.999, latency_target is 0.99, and no_data is investigate. These two targets are this lab's hypothetical policy.
One user task and one HTTP attempt are different. This task uses the attempt as the unit. Do not fill an unobserved window with a success rate of 1.
Remove the bug that counts only the success path
Implement classify(path, status, latency_ms) in classifier.py. It returns three boolean fields: eligible, available, and fast. eligible=True only when the path with the query string removed is /provision and the status code is 2xx or 5xx. available is the 2xx among them. fast may be left False in this step. Exclude /healthz, other paths, and 4xx.
503 and 504 are not successes but go into the denominator. Do not include user values from the URL query in path classification or metric labels.
Do not count a slow 201 as a good response
Complete fast in classifier.py. It is True only when available=True and latency_ms <= 1000. 1000 is included and 1000.1 is excluded. The input latency is a finite number of milliseconds of 0 or more. Keep the eligible and available contracts, and return real bools rather than strings or 0 and 1.
A fast 503 is also not a good-latency response. This SLI is the ratio of fast successful responses among all valid attempts, and it does not take only successful requests as the denominator.
Compare seven responses with four kinds of metrics
Save the output of the helper's exercise /root/cnpe-sli to http-receipt.json. The actual status codes are, in order, 200,201,201,503,504,400,404, and the server processing time of the slow 201 is 1000ms or more. It must be observed=7, eligible=4, available=2, fast=1. The result of running again with the current classifier must be the same too.
Even if the classify unit test is correct, things can still fall out on the HTTP path. Do not look only at the saved numbers; replay with the current code whether failed responses are included in the denominator.
Separate duplicate collection from retries
Implement summarize(events) in summarize.py. Each event has the fields attempt_id, path, status, and latency_ms. A duplicate with the same attempt_id and the same content is classified only once and increments duplicates. The same ID with different content is a ValueError. Retries with different IDs are each counted. The return fields are five integers: eligible, available, fast, excluded, and duplicates. Reuse classifier.classify.
If you cover a failed attempt a and a successful retry b with one final success, the attempt-based SLI changes. Conversely, when the collector sent the same event twice, you do not put it in the denominator twice.
Handle the 0.1-event budget and observation gaps
Implement decide(eligible, good, target, covered) in budget.py. It accepts only ints with 0 <= good <= eligible, a number with 0 < target < 1, and a bool covered, and otherwise raises ValueError. If covered=False or eligible=0, return sli, allowed_bad, remaining_bad, and consumed as null and decision as investigate. For the rest, sli=good/eligible, allowed_bad=eligible*(1-target), remaining_bad=allowed_bad-(eligible-good), and consumed=(eligible-good)/allowed_bad. Round sli and consumed to 6 decimal places. Do not round the budget to an integer. If the remaining budget <= 0, it is freeze, otherwise ship.
With a 99.9% target over 100 attempts, the allowed failure amount is 0.1. Do not turn it into the integer 0. Decimal(str(target)) is one way to compare the budget boundary without turning 0.999 into a binary floating-point error. Also note that bool is a subtype of int in Python.
Keep good availability from hiding bad latency
Save the helper's windows output as windows.json and implement build(windows) in report.py. The input is a list of independent windows that have name, eligible, available, fast, and covered. The return is an object keyed by window name that has availability=decide(eligible,available,0.999,covered), latency=decide(eligible,fast,0.99,covered), and decision. The combined judgment is investigate first, then freeze, and finally ship. A duplicate window name is a ValueError. Save this result to report.json.
steady is ship, slow is freeze because of latency even if availability is good, and gap is investigate even if the numbers look perfect. Not just the output file but the function too must compute correctly for new window input.
Check that your tests catch wrong implementations
In counterexamples.json, write 5–24 cases of {op,args,expected}. op is classify, summarize, or decide, and args is the list of positional arguments for that function. expected is the exact return object, and when expecting a ValueError, {"error":"ValueError"}. Include cases for all three functions, and you must catch all five errors: dropping 5xx, treating a slow success as fast, double-counting duplicates, rounding the budget to an integer, and approving an observation gap.
A test that looks only at healthy cases can let wrong implementations pass too. For each function, choose inputs where the original implementation and the wrong implementation produce different outputs. Writing wrong expected values to make a red light is not validation either.