TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Metrics Diverge in the Computation, Not the Definition

Continue in TT Lab

In one sentence

Knowing the names of the four DORA metrics is not the end of it. Even with the same data, the numbers can differ by multiples depending on how you choose the window, denominator, and representative value, so a metric's definition together with its calculation method must be pinned down in code, not in a document.

Why this was needed

There was a case where two teams looked at the same deployment records and each reported "lead time 40 minutes" and "lead time 4 hours." The data was the same. One used the median and the other the mean, and one counted only successful deployments while the other counted all of them. When this happens, you cannot have a conversation in numbers. Organizations where meetings get longer the more metrics there are are usually stuck here.

So the first thing a platform team must decide when introducing metrics is not which metrics to look at but how to calculate them.

How it works

Start from the ledger

Metrics come from a ledger, not from a dashboard. For each deployment, you leave one row with the service name, when the commit was created and when it was deployed, whether it succeeded or failed, and, if it failed, how long recovery took. With this one sheet you can calculate all four metrics, and without it, no matter which tool you buy, you only get estimates.

Where the calculation of the four metrics actually diverges

Metric Where people often go wrong
Deployment frequency If you look only at the organization-wide total, one frequently deploying service hides the rest. Look per service as well
Lead time for changes The mean gets dragged by the single case that took long. Use the median
Change failure rate The denominator is the number of deployments. Not the number of services or incidents
Mean time to restore The denominator is the number of failed deployments. If you divide by all deployments, the value gets several times smaller

There is one more trap with the median. The mean of per-service medians is not the overall median. Because the number of deployments differs per service, the two values generally differ, and if you mix them, the same ledger produces different reports.

Rate by writing the thresholds down first

DORA divides the four metrics into Elite, High, Medium, and Low. What matters here is not the thresholds themselves but writing down in advance the thresholds our organization uses and applying them as written. If you adjust the criteria a little each time you judge, the ratings lose all meaning.

And the four metrics are read together. If the speed metrics are good but the stability metrics are bad, that organization is not fast but skipping verification. Conversely, if only stability is good and deployments are infrequent, improvement may have stopped because they are avoiding risk.

For adoption rate, the denominator and the definition of 'active' are everything

The two things you always end up arguing about in platform adoption rate are whether the denominator is all services or onboarded services, and what counts as 'using it.' If you count onboarding completion as adoption, the number looks nice but hides the state where nobody is actually using it. It is more honest to count services that actually deployed within a recent period, and you must write that period down too so you can compare later.

The error budget is arithmetic

If the SLO is 99.5% and the window is 28 days, the budget comes out like this. 28 days is 40320 minutes, and 0.5% of that, 201.6 minutes, is the bad time allowed in this window. Subtract the time actually consumed and you get the remaining budget. Only with this number can you talk about "should we ship more new features now, or spend on stabilization?" in terms of remaining amount rather than emotion.

What it looks like in the field

Once you start using metrics as a ranking table between teams, the ledger gets polluted. People stop recording failures as failures and split deployment units into smaller pieces to increase the count. The numbers get better while the actual situation gets worse, and from then on the metrics become a device that hides reality.

So a platform team aims metrics at itself. If there is a rule with a low compliance rate, instead of calling in those teams, it makes that rule easier to follow, and if a service has an unusually long lead time, it looks at what is waiting in that service's pipeline. The use of a metric lies not in ranking but in deciding what to fix next.

How metrics change people's behavior

There is something you soon learn once you start measuring platform metrics. The moment you measure it, that number becomes the goal, and people move in the direction that makes the number look good. So what you measure becomes what happens.

If you measure only deployment frequency, meaningless deployments increase. Redeployments with no change and finely split commits push the number up. That is why the four metrics are looked at together. It is a real improvement only if deployment frequency rises while the change failure rate stays the same, and if only one got better, usually the other was sacrificed.

If you do not define the denominator of the change failure rate, any value can come out. Depending on whether "failure" means a rollback, an outage, or an incident at or above a severity grade, the rate differs by multiples. Write the definition down, and when you change it, recalculate past values too. Otherwise the graph of the month the definition changed looks like an improvement.

For recovery time, look at the median and the worst case together. If most take 10 minutes to finish but once a quarter one takes 8 hours, the mean tells you nothing. What people remember is those 8 hours.

Do not use it for comparison between teams. If the nature of the services differs, the numbers differ too. If you put a payment system and an internal tool in the same table, the payment team moves toward not taking risks. Metrics are for looking at change over time within the same team.

Define adoption rate not as "using it" but as "working with it." The number of people who created an account says nothing. The definition must be one that counts behavior, such as the number of teams that actually deployed in the last 30 days, for you to know whether the platform is helping.

The error budget is a negotiation tool. It has value only when there is an agreement in advance that if budget remains you may make risky changes, and when it is used up you focus on stabilization. Without that agreement, measuring only the number just turns it into material to argue over after an outage.

What to do in the next lab

With a 28-day deployment ledger, you calculate the four metrics yourself. You check by hand the difference between median and mean and the denominators of failure rate and recovery time, and rate them with predetermined thresholds. Next, you calculate the adoption rate from the service roster according to the definition of active use, and finally work out the error budget's budget and burn rate.