TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

Should you deploy with an error budget of 0.1 requests?

Continue in TT Lab

One-line summary

An error budget is not dashboard decoration but an input for deciding whether to change, based on verified observation and an agreed policy. If you approve a deployment because there are 0 errors even though there is no data, that is not a safe judgment but a judgment without evidence.

Why this was needed

A hypothetical platform team set 99.9% issuance availability and a 99% fast-success ratio for 30 days. In this window there are 10000 valid attempts, 9995 available successes, and 9850 fast successes. Looking at availability alone, it is 99.95%, above the target, but the fast successes are 98.5%, which falls short. If the team deploys a new feature looking at only the availability cell, it can make an already bad latency experience worse.

The figures in this unit are examples for learning policy judgment. They do not mean the right target for every platform. A target that is too loose, and an extreme target that users cannot distinguish, can both distort cost and priorities. You must be able to explain which user experience you are trying to protect in the first place.

How it works

The four numbers of a request-based budget

In the same observation window, N is the number of valid attempts, G is the number of good attempts, and T is the target ratio. The allowed failure amount is N × (1 − T), the actual failure amount is N − G, and the remaining budget is the allowed failure amount minus the actual failure amount. The consumption ratio is the actual failure amount divided by the allowed failure amount. The lab displays sli and consumed to 6 decimal places, but it does not round the allowed failure amount to an integer for comparing the budget boundary.

With N=10000 and T=0.999, the allowed failures are 10. If there are 5 failures, 5 remain and half is consumed. With 10 failures, the remainder is exactly 0. With 20 failures, the remainder is -10 and the consumption ratio is 2, showing over-use. If you clip the negative to 0 to make the screen look neat, you lose the amount of the overage.

If you apply the same target when N=100, the allowed failure amount is 0.1. This does not mean there are 0.1 actual failure events; it is the arithmetic amount the ratio target allows. With 1 failure, you use 10 times the budget of this small window. If you first round 0.1 to 0, the denominator becomes 0 and even the policy interpretation goes wrong. You must explain the sensitivity of a small sample and the representative observation window together.

Budget exhaustion and the change policy are agreed

The Google SRE error budget policy example is a reference that connects budget state to change priorities and documents responsibilities and exceptions. The simple policy in this lab is freeze for general feature changes if the remaining budget is 0 or below, and ship if any remains. Do not read freeze as "all work is forbidden." Outage recovery and urgent security fixes need a separate judgment and review procedure. The function here does not call a real deployment system and only returns a learning conclusion.

If there is an observation gap or there are 0 valid attempts, it is investigate. In this case you leave the success rate and budget numbers as null to show there is no basis to compute. You must investigate next whether there really were no requests in the window, whether the collector stopped, or whether routing changed. The input covered is a flag that summarizes evidence about the observation range, and the fact that the code wrote it as true does not prove actual observation completeness.

Do not mix numbers from different windows

The steady, slow, and gap in this lab are not three consecutive days of the same service but hypothetical cases of independent 30-day observation windows. Within each window you match the denominator and the numerator. You do not divide today's failures by yesterday's denominator, nor directly call the 30-day total budget and the failure count over 5 minutes the same consumption rate. Designing alerts that use the consumption speed over a short period needs a separate time range and ratio definition.

When you combine availability and latency judgments, investigate takes priority, then freeze, and finally ship. This is so that you do not approve just because another metric is good when one metric has no evidence. The priority is this task's contract, and a more complex service policy can also include service criticality, dependencies, and change risk.

What it looks like in the field

steady has available successes of 9995/10000 and fast successes of 9950/10000, and the observation is complete, so both budgets remain. slow has the same availability but only 9850 fast successes, so it is freeze. gap is investigate because covered=False even if the collected numbers are 100%. You do not submit the three cases as only a single status string; you leave each SLI, allowed amount, remaining amount, and consumption ratio together so that the next person can recompute the grounds for the judgment.

This kind of judgment is a job competency beyond installing tools. The Supabase SRE job posting checked on 2026-09-11 links user-experience-based SLI/SLO, error budget policy, and operational readiness review. This lab covers a small compute-and-verify loop among them, and it does not mean you have met large-scale operations experience or all of the hiring requirements.

What to do in the next lab

First you classify so that HTTP failures are not dropped from the denominator and verify with real requests. After that you implement aggregation, the budget, and the report, and create healthy cases and counterexamples together. The last step checks whether your cases distinguish five wrong implementations. You should not make it fail by writing wrong expected values; it must pass the correct implementation and reject only the wrong ones.