CNPE — Cloud Native Platform Engineer
What Does a Platform Measure Itself By
One-line summary
What the platform team should measure is not the CPU of the nodes but the promises the platform has made to its users. Those promises are the success rate of self-service requests, the time it takes for a deployment to finish, and the share of deployments that had to be rolled back.
Why resource metrics are not enough
A dashboard that looks only at node metrics stays green even when an incident happens. Even if all self-service issuance fails, node CPU is actually idle. Conversely, even if nodes are busy, there may be no problem at all from the user's point of view.
So you treat the platform as a service too. The users are development teams, the requests are claim creations or deployments, and there are successes and failures and there is elapsed time. Framed this way, what to measure becomes clear.
| What to measure | Why that |
|---|---|
| The success ratio of issuance requests | Whether self-service is really self-service |
| Time from request to usable | Whether the golden path is really fast |
| Deployment frequency and the rolled-back ratio | Whether deployment is something to fear or not |
| Time taken to recover from an incident | The time to learn of the problem and the time to fix it |
How it works
The recording rule comes first, and the alert comes next
If you compute a ratio inside the alert condition every time, the same expression gets scattered across several alerts. This is where the incident of fixing one place and forgetting another comes from. So you define the ratio once as a recording rule, and the alerts reference that name.
groups:
- name: platform-slo
rules:
- record: platform:provision_success:ratio5m
expr: |
sum(rate(platform_provision_total{result="success"}[5m]))
/
sum(rate(platform_provision_total[5m]))
If the name contains 5m, the range of the expression must also be 5 minutes. If the name and the content are out of step, everyone who uses that metric judges from a wrong premise.
Choose the firing delay to fit the purpose
for requires the condition to hold for a period of time. This lab uses 10 minutes, but for an event that needs an immediate response you may omit it. This time is different from Alertmanager's send and repeat intervals. An alert expression must remove the series when it is false. vector(0) and 비율 < bool 0.95 (where the Korean word stands for the ratio) are traps in which the series remains even if the value is 0 and the alert becomes active. In the official alerting rules explanation, read while distinguishing pending from firing.
An alert without a runbook_url is the same. If an alert you receive at three in the morning does not say what to do next, it is not information but noise. The severity label is the basis on which routing splits, so without it all alerts go down the same path.
What you verified and what you deployed must be the same
A rules file can be verified with promtool check rules. A person reading it cannot catch a single misplaced parenthesis in an expression. But if the file you verified and the PrometheusRule uploaded to the cluster are different things, that verification guarantees nothing. You create the CR from the file, and after uploading it, compare the whole spec.groups. If only the name is the same and the expression changed, it is a different rule. Do not assert that the Operator selected it or that Prometheus loaded it just because the API stored it. This environment does not run the Operator, so it verifies up to storage agreement and does not claim that alert delivery was tested.
What it looks like in the field
Consider a hypothetical case. If you built an issuance service whose failure path leaves no metric, the dashboard may show only a 100% success rate. If the denominator counts only successes, the ratio is always 1.
With a single ratio alone, it is hard to tell such an instrumentation gap. It shows only when you read directly from the expression what is being counted. So when you create a new metric, you must cause a failure on purpose once and confirm that the value moves. The same goes for alerts.
Put events behind the syntax
The simple ratio above is the starting point for when both result series exist. If only the failure series exists, the numerator is empty, and if there is no traffic, the denominator is 0. In the lab that follows, when there are only failures, you supplement the numerator with 0 but leave the ratio only when the total rate is positive. An observation gap must be investigated as a separate instrumentation anomaly, not as 100% success. If you simply average the per-instance success rates, you get the error of counting a server with 1 request and a server with 10,000 requests with the same weight.
Following the official rate explanation, you first compute the rate per counter and then combine. If you build the sum first and then compute the rate, one series' restart gets mixed with another series' increase and the reset interpretation changes. The lab also checks, within the reset interval, an event in which only the success counter restarts.
The 13 events used in this process are healthy 99%, sustained failure 80%, the 95% boundary and both sides of it, a 3-minute outage, recovery after an outage, no traffic, no observation, only successes, only failures, a counter reset, and two instances with different request volumes. It also checks the success rate at 17 minutes of recovery and the resolution at 19 minutes, to tell apart the error of changing a 5-minute range to 10 minutes. This is a learning test on fixed data and not proof for all production traffic.
promtool check rules checks syntax, and promtool test rules checks values on virtual time series. Read input_series and eval_time in the official test format and distinguish them from real waiting time. The image's 3.0.1 does not support fuzzy_compare from the latest documentation, so the helper directly checks that the ratio error is below 1e-9.
What to read next
After a metric tells you about an anomaly, you next look at what to count first to classify that anomaly.