TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

What Must You Measure for the Platform to Improve

Continue in TT Lab

In one sentence

The four DORA metrics are tools for measuring the health of the delivery pipeline, not for measuring the success of a platform. A platform must be measured separately by adoption rate, lead time, and cognitive load, and the platform itself should also have an SLO.

Why this was needed

The most common mistake when a platform team reports results is to list "what we built": "12 CI templates, 30 charts, 40 dashboards." That is output, not outcome. It may be that nobody is using those 12 templates.

To measure outcomes, you have to look from the user's side. And the users are other development teams.

How it works

The four DORA metrics and their limits

Metric Meaning
Deployment Frequency How often you ship to production
Lead Time for Changes Time for a commit to reach production
Change Failure Rate The share of deployments that led to an incident
Time to restore service (MTTR / Failed Deployment Recovery Time) Time to recover from a failure

The first two are speed (throughput), and the last two are stability. The key finding of the DORA research was that the two do not conflict — fast organizations are generally more stable.

The limits are also clear.

So use DORA as an alarm device, and measure the platform's outcomes separately.

Platform-specific metrics

A trap that often appears here: running the survey only once is useless. Absolute values cannot be interpreted; only the trend is meaningful.

The platform should have an SLO too

The moment a platform becomes something other teams depend on, the platform's availability is their availability. Yet many platform teams operate without an SLO.

Examples of platform SLOs:

With an SLO you get an error budget, and with an error budget you can decide "whether to ship new features or stabilize this quarter" by numbers rather than gut feeling. It is the platform team applying SRE methodology to itself.

Multi-tenancy isolation levels

A platform makes multiple teams share resources. The options for isolation level are roughly these.

Level Isolation Cost Operational burden
Namespace Logical. Separated by RBAC, quotas, and NetworkPolicy Cheapest Low
Node pool separation Per-tenant nodes. Placement through taint/toleration Medium (idle resources arise) Medium
Cluster separation Separated down to the control plane Expensive High (upgrades for as many clusters as you have)
Account or subscription separation Separated down to the cloud boundary Most expensive Highest

The selection criteria are "what is the trust relationship between these tenants?" and "is there a regulatory requirement?" For teams in the same company, namespaces are often enough, and if external customers run code, you need at least node separation and usually cluster separation. The vocabulary that distinguishes soft multi-tenancy (trusted tenants) from hard multi-tenancy (untrusted tenants) appears on the exam.

FinOps — cost is developer experience too

Cost should be a feedback signal, not an after-the-fact bill. There are three core practices.

  1. Attribution (showback/chargeback) — map costs to teams through labels and namespaces. If nobody knows whose it is, nobody reduces it. Showback only shows, while chargeback actually bills.
  2. The gap between requests and actual usage — the biggest waste in Kubernetes cost is excessive requests. Just showing teams the actual usage rate against requests reduces a large part of it.
  3. Put right-sizing into the golden path — if the request values in the default template are reasonable, costs are saved even when nobody pays attention.

What it looks like in the field

The author's homelab has observability in place with kube-prometheus-stack and Grafana (10.0.0.203), and the list of remaining tasks includes "observe control plane load — how much etcd + apiserver a 4C/14GB mini PC can take." This is a miniature of platform SLO thinking. When the control plane slows down, every team on top of it slows down, so the platform's resource headroom is not only the platform's problem but the whole set of users' problem.

The isolation story also appears as is in the same cluster. The GPUs are 24GB, 32GB, and two 8GB cards, each in a different weight class, so with no measures in place a large training job lands on a small card. This is a performance problem and also a fairness problem — if one tenant monopolizes the large card, other tenants' lead time grows. So a placement rule was made with the gpu.homelab/tier label and nodeSelector. Isolation is a topic not only of security but of predictability.

And this cluster's repeated lesson that "status is Ready" and "it actually works" are different claims applies to measurement as well. Even when every component health check is green, the experience users have can be bad. Platform metrics must be measured on the user journey, not on component status.

What to check in the next quiz

This module ends with a quiz. It is worth revisiting how the CRD and quotas built in the earlier module and the golden path scaffold connect to the metrics discussed here — for example, the fact that the defaults of quotas and LimitRange directly affect cost metrics.