CNPA — Cloud Native Platform Engineering Associate
Computing Platform Metrics From the Deployment Ledger
Goal
From a single 28-day deployment ledger, you calculate the four DORA metrics, the adoption rate, and the error budget yourself. You build the habit of extracting values from the ledger rather than writing them by eye.
Why it matters
Most arguments over metrics start not from the definition but from the calculation method. Depending on whether it is the mean or the median, whether the denominator is the number of deployments or the number of failures, and what counts as active use, the same data yields numbers that differ by multiples. The first thing a platform team should do when introducing metrics is to pin down the calculation in code so that anyone who runs it gets the same value. The grader for this lab works on the same principle. It does not trust the numbers you wrote and recalculates from the ledger to compare.
Steps
- Copy the deployment ledger into
/root/cnpa-metrics/deploys.csv. The header isservice,day,lead_min,result,restore_minand there are 15 data rows. The contents are exactly as in the format example.resultissuccessorfailed, andrestore_minof a successful deployment is 0. - Put the deployment frequency into
/root/cnpa-metrics/freq.json. The keys arewindow_days(28),weeks(4), andper_week, and inper_weekyou put each service's weekly deployment count to two decimal places. Include only the three services in the ledger. - Put the lead time medians into
/root/cnpa-metrics/lead.json. The keys aremedian_min(the per-service medians) andoverall_median_min(the median of all 15 records). - Put the change failure rate into
/root/cnpa-metrics/cfr.json. The keys aretotal,failed, andcfr_pct(to one decimal place). - Put the recovery time into
/root/cnpa-metrics/mttr.json. The keys arefailed,mean_restore_min(the mean recovery time of failed deployments, to one decimal place), andmax_restore_min. - Put the ratings of the four metrics into
/root/cnpa-metrics/dora.json. The keys aredeploy_frequency,lead_time,change_failure_rate, andmttr, and each value is one ofElite,High,Medium, andLow. Use these thresholds. For deployment frequency (weekly count, dividing all deployments by 4 weeks), 7 or more is Elite, 1 or more is High, 0.25 or more is Medium, and below that is Low. For lead time (overall median, in minutes), below 60 is Elite, below 1440 is High, below 10080 is Medium, and anything above is Low. For change failure rate, 5% or less is Elite, 10% or less is High, 15% or less is Medium, and anything larger is Low. For recovery time (mean, in minutes), use the same thresholds as lead time. - Copy the service roster into
/root/cnpa-metrics/services.csv. The header isservice,onboarded_day,golden_path,last_deploy_dayand there are 6 data rows. Then put the adoption status into/root/cnpa-metrics/platform-adoption.json. The keys aretotal_services(the number of services in the roster),onboarded(the count wheregolden_pathisyes),active(the count wheregolden_pathisyesandlast_deploy_dayis 15 or more), andactive_rate_pct(the active count divided by the total number of services, as a percentage, to one decimal place). - Put the error budget into
/root/cnpa-metrics/budget.json. The platform SLO is 99.5% and the window is 28 days. The keys areslo_pct(99.5),window_days(28),budget_min(0.5% of the window's total minutes),burned_min(the sum of recovery times in the ledger),burn_pct(the burn rate, to one decimal place), andremaining_min(the remaining budget).
Notes
- For the median, sort with
sort -nand pick the middle value. With an odd count, it is the single middle one. - Set decimal places with
awk 'BEGIN { printf "%.1f", ... }'. - If you build JSON with
jq -n --argjson, numbers are not turned into strings. - The grader compares your JSON against the ledger. If you try to fix the ledger to make the numbers match, step 1 grading fails first.
Copy the deployment ledger
Copy the deployment ledger into /root/cnpa-metrics/deploys.csv. The header is service,day,lead_min,result,restore_min and there are 15 data rows. The contents are exactly as in the format example. result is success or failed, and restore_min of a successful deployment is 0.
This is the starting point of every calculation. If even one character differs, all the values in later steps differ, so check with totals after copying.
Deployment frequency
Put the deployment frequency into /root/cnpa-metrics/freq.json. The keys are window_days (28), weeks (4), and per_week, and in per_week you put each service's weekly deployment count to two decimal places. Include only the three services in the ledger.
Count each service separately. 28 days is 4 weeks, and write values to two decimal places.
Lead time median
Put the lead time medians into /root/cnpa-metrics/lead.json. The keys are median_min (the per-service medians) and overall_median_min (the median of all 15 records).
It is the middle value after sorting. The overall median is found by laying out all 15 records. It is not the mean of the per-service medians.
Change failure rate
Put the change failure rate into /root/cnpa-metrics/cfr.json. The keys are total, failed, and cfr_pct (to one decimal place).
The denominator is the number of deployments. Not the number of services, nor the number of incidents. Write the percentage to one decimal place.
Recovery time
Put the recovery time into /root/cnpa-metrics/mttr.json. The keys are failed, mean_restore_min (the mean recovery time of failed deployments, to one decimal place), and max_restore_min.
The number you divide by when taking the mean is the number of failed deployments. If you divide by all deployments, the value gets several times smaller.
Rate the metrics
Put the ratings of the four metrics into /root/cnpa-metrics/dora.json. The keys are deploy_frequency, lead_time, change_failure_rate, and mttr, and each value is one of Elite, High, Medium, and Low. Use these thresholds. For deployment frequency (weekly count, dividing all deployments by 4 weeks), 7 or more is Elite, 1 or more is High, 0.25 or more is Medium, and below that is Low. For lead time (overall median, in minutes), below 60 is Elite, below 1440 is High, below 10080 is Medium, and anything above is Low. For change failure rate, 5% or less is Elite, 10% or less is High, 15% or less is Medium, and anything larger is Low. For recovery time (mean, in minutes), use the same thresholds as lead time.
Use the thresholds exactly as written in the instructions. Deployment frequency is judged by the organization-wide total (15 records) divided by 4 weeks.
Adoption rate
Copy the service roster into /root/cnpa-metrics/services.csv. The header is service,onboarded_day,golden_path,last_deploy_day and there are 6 data rows. Then put the adoption status into /root/cnpa-metrics/platform-adoption.json. The keys are total_services (the number of services in the roster), onboarded (the count where golden_path is yes), active (the count where golden_path is yes and last_deploy_day is 15 or more), and active_rate_pct (the active count divided by the total number of services, as a percentage, to one decimal place).
Onboarding and active use are different. Active means a service onboarded through the golden path that deployed within the last 14 days. The denominator of the active rate is all services.
Error budget
Put the error budget into /root/cnpa-metrics/budget.json. The platform SLO is 99.5% and the window is 28 days. The keys are slo_pct (99.5), window_days (28), budget_min (0.5% of the window's total minutes), burned_min (the sum of recovery times in the ledger), burn_pct (the burn rate, to one decimal place), and remaining_min (the remaining budget).
28 days is 40320 minutes. Since the SLO is 99.5%, the 0.5% of that is this window's budget, and the burn is the sum of recovery times in the ledger.