CNPA — Cloud Native Platform Engineering Associate
What Must You Measure for the Platform to Improve
In one sentence
The four DORA metrics are tools for measuring the health of the delivery pipeline, not for measuring the success of a platform. A platform must be measured separately by adoption rate, lead time, and cognitive load, and the platform itself should also have an SLO.
Why this was needed
The most common mistake when a platform team reports results is to list "what we built": "12 CI templates, 30 charts, 40 dashboards." That is output, not outcome. It may be that nobody is using those 12 templates.
To measure outcomes, you have to look from the user's side. And the users are other development teams.
How it works
The four DORA metrics and their limits
| Metric | Meaning |
|---|---|
| Deployment Frequency | How often you ship to production |
| Lead Time for Changes | Time for a commit to reach production |
| Change Failure Rate | The share of deployments that led to an incident |
| Time to restore service (MTTR / Failed Deployment Recovery Time) | Time to recover from a failure |
The first two are speed (throughput), and the last two are stability. The key finding of the DORA research was that the two do not conflict — fast organizations are generally more stable.
The limits are also clear.
- They are easy to game. If you split deployments into smaller pieces, frequency goes up. If you narrow the definition of failure, the failure rate goes down.
- They erase per-team context. Comparing a payment service that requires regulatory review with an internal dashboard using the same yardstick leads to wrong conclusions.
- They cannot directly show the platform's contribution. When lead time drops, the metrics alone cannot tell you whether it is thanks to the platform or because the team got smaller.
- They can diverge from the pain developers actually experience. In an organization where deployment is fast but setting up a local development environment takes two days, DORA is all green.
So use DORA as an alarm device, and measure the platform's outcomes separately.
Platform-specific metrics
- Adoption — how much it is used when not mandated. The number of actively using teams, and the share of services created through the golden path.
- Time to first deploy — the time for a new team or new service to get from nothing to a production deployment. It measures the platform's core promise most directly.
- Number and types of support tickets — if they decrease, it means self-service is taking hold. The types that increase are the next golden path candidates.
- Cognitive load survey — a qualitative metric, but the only one that directly measures the goal. You ask questions such as "How easy was it to find what you needed?" and "How many days do you think this task would take?" in the same form every quarter and watch the trend.
A trap that often appears here: running the survey only once is useless. Absolute values cannot be interpreted; only the trend is meaningful.
The platform should have an SLO too
The moment a platform becomes something other teams depend on, the platform's availability is their availability. Yet many platform teams operate without an SLO.
Examples of platform SLOs:
- Deployment pipeline success rate and P95 duration
- Completion time of self-service provisioning requests (for example, 95% of new namespaces within 2 minutes)
- Availability of the internal API and portal
- Success rate of golden path scaffolding
With an SLO you get an error budget, and with an error budget you can decide "whether to ship new features or stabilize this quarter" by numbers rather than gut feeling. It is the platform team applying SRE methodology to itself.
Multi-tenancy isolation levels
A platform makes multiple teams share resources. The options for isolation level are roughly these.
| Level | Isolation | Cost | Operational burden |
|---|---|---|---|
| Namespace | Logical. Separated by RBAC, quotas, and NetworkPolicy | Cheapest | Low |
| Node pool separation | Per-tenant nodes. Placement through taint/toleration | Medium (idle resources arise) | Medium |
| Cluster separation | Separated down to the control plane | Expensive | High (upgrades for as many clusters as you have) |
| Account or subscription separation | Separated down to the cloud boundary | Most expensive | Highest |
The selection criteria are "what is the trust relationship between these tenants?" and "is there a regulatory requirement?" For teams in the same company, namespaces are often enough, and if external customers run code, you need at least node separation and usually cluster separation. The vocabulary that distinguishes soft multi-tenancy (trusted tenants) from hard multi-tenancy (untrusted tenants) appears on the exam.
FinOps — cost is developer experience too
Cost should be a feedback signal, not an after-the-fact bill. There are three core practices.
- Attribution (showback/chargeback) — map costs to teams through labels and namespaces. If nobody knows whose it is, nobody reduces it. Showback only shows, while chargeback actually bills.
- The gap between requests and actual usage — the biggest waste in Kubernetes cost is excessive requests. Just showing teams the actual usage rate against requests reduces a large part of it.
- Put right-sizing into the golden path — if the request values in the default template are reasonable, costs are saved even when nobody pays attention.
What it looks like in the field
The author's homelab has observability in place with kube-prometheus-stack and Grafana (10.0.0.203), and the list of remaining tasks includes "observe control plane load — how much etcd + apiserver a 4C/14GB mini PC can take." This is a miniature of platform SLO thinking. When the control plane slows down, every team on top of it slows down, so the platform's resource headroom is not only the platform's problem but the whole set of users' problem.
The isolation story also appears as is in the same cluster. The GPUs are 24GB, 32GB, and two 8GB cards, each in a different weight class, so with no measures in place a large training job lands on a small card. This is a performance problem and also a fairness problem — if one tenant monopolizes the large card, other tenants' lead time grows. So a placement rule was made with the gpu.homelab/tier label and nodeSelector. Isolation is a topic not only of security but of predictability.
And this cluster's repeated lesson that "status is Ready" and "it actually works" are different claims applies to measurement as well. Even when every component health check is green, the experience users have can be bad. Platform metrics must be measured on the user journey, not on component status.
What to check in the next quiz
This module ends with a quiz. It is worth revisiting how the CRD and quotas built in the earlier module and the golden path scaffold connect to the metrics discussed here — for example, the fact that the defaults of quotas and LimitRange directly affect cost metrics.