TT Lab
Get started
Learn Learning paths Courses

Observability

Computing the Error Budget as a Number

Continue in TT Lab

Goal

You observe the recent error rate on a real Prometheus and build a feature-deployment decision program using a separate synthetic 30-day aggregate. You distinguish burn rate from period usage, and freeze from insufficient data.

Why it matters

The 12-hour error rate divided by the allowed error rate is the 12-hour burn rate. You cannot conclude from that one value that the 30-day budget has already been spent. Substituting 0 when there is no observation is not a safe decision either. After checking the units and scope of the real data, you execute an explicit policy as code.

The estimated time is 75 minutes. Extend it with the + hour button before the default 60-minute session ends (up to 180 minutes). When the session ends, the files in /root/obs disappear, so keep the code you need elsewhere. Python conditionals, dictionaries, and JSON input and output are prerequisite knowledge.

Steps

  1. Write the success ratio SLI of the synthetic service, where only 5xx is defined as failure, to /root/obs/slo-01-sli.promql. Sum the rate over the last 1 hour to get successful requests / total requests. Do not generalize this definition to 4xx of other services.
  2. Calculate the 30-day allowed downtime of a time-based 99.9% SLO in minutes and write only the number to /root/obs/slo-02-budget.txt. Do not convert it into a request error count.
  3. Write the burn rate query, the error rate of the last 12 hours divided by 0.001, to /root/obs/slo-03-consumed.promql. The file name is kept for compatibility with the existing lab, but the value is not monthly usage.
  4. Write the burn rate query for the last 1 hour to /root/obs/slo-04-burn1h.promql. The window length differs from step 3, and both are speed metrics.
  5. Write to /root/obs/slo-05-multi.promql the condition that the 1-hour and 5-minute burn rates each exceed 14.4, joined with and. Do not ignore a short window that has recovered.
  6. Write a multi-window alert rule named ErrorBudgetBurnFast in /etc/prometheus/rules/slo.yml. Include alert, expr, for, labels.severity, and annotations.summary, and check it with promtool check rules. Alert delivery and the feature deployment allow policy are separate matters.
  7. After reloading Prometheus, confirm through the rules API that the alert is registered. The fact that it is registered alone does not verify actual firing or operational response.
  8. Read /opt/lab/slo_release/contract.md and complete /root/obs/slo-gate.py. From one synthetic observation JSON on stdin, compute the period budget usage ratio and the burn rates of the two windows. Insufficient data is hold, a budget usage ratio of 1 or more is freeze, apart from that if both windows exceed 14.4 it is freeze, and the rest is allow. The output has five fields: decision, reason, budget_used, burn_hour, and burn_five_minutes, and the exit codes are allow=0, freeze=2, hold=3. The order of data checks and the detailed reasons follow the execution contract.

Notes

Define the SLI as a query

Success ratio SLI → /root/obs/slo-01-sli.promql

This synthetic service counts only observed 5xx as failure. Divide the sum of the rate of successful requests by the sum of the total rate. Do not unconditionally classify 4xx of other services as success.

The 30-day error budget in minutes

Time-based 30-day budget (minutes) → /root/obs/slo-02-budget.txt

This is a time-based 99.9% SLO. Multiply the number of minutes in 30 days by the allowed downtime ratio and write it to one decimal place. Do not confuse it with a request-count budget.

Observe the 12-hour burn rate

Observe the 12-hour burn rate → /root/obs/slo-03-consumed.promql

Divide the 12-hour error rate by the allowed error ratio 0.001. Exceeding 1 means the speed in that window is faster than the allowed speed, not that the 30-day budget has already been used up.

Observe the 1-hour burn rate

Observe the 1-hour burn rate → /root/obs/slo-04-burn1h.promql

Change the window to 1h with the same error definition. Both the long window and the short window are burn rates, and you distinguish them from budget usage.

Multi-window condition

Multi-window condition → /root/obs/slo-05-multi.promql

Join with and the conditions that 1 hour and 5 minutes each exceed 14.4. You must remove false elements. Even a 0 from a bool comparison can turn the alert on. Use the time-series checks in /opt/lab/slo_release/promql.md to confirm recovery and missing data.

Write the alert rule

ErrorBudgetBurnFast multi-window alert rule → /etc/prometheus/rules/slo.yml

Include alert, expr, for, labels.severity, and annotations.summary. for: 0s evaluates the two-window condition with no additional wait for persistence. A long for can delay even the alert for a serious failure.

Confirm the rule is registered

Confirm that the ErrorBudgetBurnFast rule is registered in Prometheus

After the configuration check, reload and confirm the name in the rules API. Even if it is inactive, the registration check passes, but it does not replace a test of actual firing and alert delivery.

Verify the deployment decision as code

Decision program → /root/obs/slo-gate.py

Implement the data checks of the execution contract first. Leave a monthly budget overrun, a recent outage in the two windows, and unconfirmed data with different reasons and exit codes.