TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Computing Platform Metrics From the Deployment Ledger

Continue in TT Lab

Goal

From a single 28-day deployment ledger, you calculate the four DORA metrics, the adoption rate, and the error budget yourself. You build the habit of extracting values from the ledger rather than writing them by eye.

Why it matters

Most arguments over metrics start not from the definition but from the calculation method. Depending on whether it is the mean or the median, whether the denominator is the number of deployments or the number of failures, and what counts as active use, the same data yields numbers that differ by multiples. The first thing a platform team should do when introducing metrics is to pin down the calculation in code so that anyone who runs it gets the same value. The grader for this lab works on the same principle. It does not trust the numbers you wrote and recalculates from the ledger to compare.

Steps

  1. Copy the deployment ledger into /root/cnpa-metrics/deploys.csv. The header is service,day,lead_min,result,restore_min and there are 15 data rows. The contents are exactly as in the format example. result is success or failed, and restore_min of a successful deployment is 0.
  2. Put the deployment frequency into /root/cnpa-metrics/freq.json. The keys are window_days (28), weeks (4), and per_week, and in per_week you put each service's weekly deployment count to two decimal places. Include only the three services in the ledger.
  3. Put the lead time medians into /root/cnpa-metrics/lead.json. The keys are median_min (the per-service medians) and overall_median_min (the median of all 15 records).
  4. Put the change failure rate into /root/cnpa-metrics/cfr.json. The keys are total, failed, and cfr_pct (to one decimal place).
  5. Put the recovery time into /root/cnpa-metrics/mttr.json. The keys are failed, mean_restore_min (the mean recovery time of failed deployments, to one decimal place), and max_restore_min.
  6. Put the ratings of the four metrics into /root/cnpa-metrics/dora.json. The keys are deploy_frequency, lead_time, change_failure_rate, and mttr, and each value is one of Elite, High, Medium, and Low. Use these thresholds. For deployment frequency (weekly count, dividing all deployments by 4 weeks), 7 or more is Elite, 1 or more is High, 0.25 or more is Medium, and below that is Low. For lead time (overall median, in minutes), below 60 is Elite, below 1440 is High, below 10080 is Medium, and anything above is Low. For change failure rate, 5% or less is Elite, 10% or less is High, 15% or less is Medium, and anything larger is Low. For recovery time (mean, in minutes), use the same thresholds as lead time.
  7. Copy the service roster into /root/cnpa-metrics/services.csv. The header is service,onboarded_day,golden_path,last_deploy_day and there are 6 data rows. Then put the adoption status into /root/cnpa-metrics/platform-adoption.json. The keys are total_services (the number of services in the roster), onboarded (the count where golden_path is yes), active (the count where golden_path is yes and last_deploy_day is 15 or more), and active_rate_pct (the active count divided by the total number of services, as a percentage, to one decimal place).
  8. Put the error budget into /root/cnpa-metrics/budget.json. The platform SLO is 99.5% and the window is 28 days. The keys are slo_pct (99.5), window_days (28), budget_min (0.5% of the window's total minutes), burned_min (the sum of recovery times in the ledger), burn_pct (the burn rate, to one decimal place), and remaining_min (the remaining budget).

Notes

Copy the deployment ledger

Copy the deployment ledger into /root/cnpa-metrics/deploys.csv. The header is service,day,lead_min,result,restore_min and there are 15 data rows. The contents are exactly as in the format example. result is success or failed, and restore_min of a successful deployment is 0.

This is the starting point of every calculation. If even one character differs, all the values in later steps differ, so check with totals after copying.

Deployment frequency

Put the deployment frequency into /root/cnpa-metrics/freq.json. The keys are window_days (28), weeks (4), and per_week, and in per_week you put each service's weekly deployment count to two decimal places. Include only the three services in the ledger.

Count each service separately. 28 days is 4 weeks, and write values to two decimal places.

Lead time median

Put the lead time medians into /root/cnpa-metrics/lead.json. The keys are median_min (the per-service medians) and overall_median_min (the median of all 15 records).

It is the middle value after sorting. The overall median is found by laying out all 15 records. It is not the mean of the per-service medians.

Change failure rate

Put the change failure rate into /root/cnpa-metrics/cfr.json. The keys are total, failed, and cfr_pct (to one decimal place).

The denominator is the number of deployments. Not the number of services, nor the number of incidents. Write the percentage to one decimal place.

Recovery time

Put the recovery time into /root/cnpa-metrics/mttr.json. The keys are failed, mean_restore_min (the mean recovery time of failed deployments, to one decimal place), and max_restore_min.

The number you divide by when taking the mean is the number of failed deployments. If you divide by all deployments, the value gets several times smaller.

Rate the metrics

Put the ratings of the four metrics into /root/cnpa-metrics/dora.json. The keys are deploy_frequency, lead_time, change_failure_rate, and mttr, and each value is one of Elite, High, Medium, and Low. Use these thresholds. For deployment frequency (weekly count, dividing all deployments by 4 weeks), 7 or more is Elite, 1 or more is High, 0.25 or more is Medium, and below that is Low. For lead time (overall median, in minutes), below 60 is Elite, below 1440 is High, below 10080 is Medium, and anything above is Low. For change failure rate, 5% or less is Elite, 10% or less is High, 15% or less is Medium, and anything larger is Low. For recovery time (mean, in minutes), use the same thresholds as lead time.

Use the thresholds exactly as written in the instructions. Deployment frequency is judged by the organization-wide total (15 records) divided by 4 weeks.

Adoption rate

Copy the service roster into /root/cnpa-metrics/services.csv. The header is service,onboarded_day,golden_path,last_deploy_day and there are 6 data rows. Then put the adoption status into /root/cnpa-metrics/platform-adoption.json. The keys are total_services (the number of services in the roster), onboarded (the count where golden_path is yes), active (the count where golden_path is yes and last_deploy_day is 15 or more), and active_rate_pct (the active count divided by the total number of services, as a percentage, to one decimal place).

Onboarding and active use are different. Active means a service onboarded through the golden path that deployed within the last 14 days. The denominator of the active rate is all services.

Error budget

Put the error budget into /root/cnpa-metrics/budget.json. The platform SLO is 99.5% and the window is 28 days. The keys are slo_pct (99.5), window_days (28), budget_min (0.5% of the window's total minutes), burned_min (the sum of recovery times in the ledger), burn_pct (the burn rate, to one decimal place), and remaining_min (the remaining budget).

28 days is 40320 minutes. Since the SLO is 99.5%, the 0.5% of that is this window's budget, and the burn is the sum of recovery times in the ledger.