Computing the Error Budget as a Number
Goal
You observe the recent error rate on a real Prometheus and build a feature-deployment decision program using a separate synthetic 30-day aggregate. You distinguish burn rate from period usage, and freeze from insufficient data.
Why it matters
The 12-hour error rate divided by the allowed error rate is the 12-hour burn rate. You cannot conclude from that one value that the 30-day budget has already been spent. Substituting 0 when there is no observation is not a safe decision either. After checking the units and scope of the real data, you execute an explicit policy as code.
The estimated time is 75 minutes. Extend it with the + hour button before the default 60-minute session ends (up to 180 minutes). When the session ends, the files in /root/obs disappear, so keep the code you need elsewhere. Python conditionals, dictionaries, and JSON input and output are prerequisite knowledge.
Steps
- Write the success ratio SLI of the synthetic service, where only 5xx is defined as failure, to /root/obs/slo-01-sli.promql. Sum the rate over the last 1 hour to get successful requests / total requests. Do not generalize this definition to 4xx of other services.
- Calculate the 30-day allowed downtime of a time-based 99.9% SLO in minutes and write only the number to /root/obs/slo-02-budget.txt. Do not convert it into a request error count.
- Write the burn rate query, the error rate of the last 12 hours divided by 0.001, to /root/obs/slo-03-consumed.promql. The file name is kept for compatibility with the existing lab, but the value is not monthly usage.
- Write the burn rate query for the last 1 hour to /root/obs/slo-04-burn1h.promql. The window length differs from step 3, and both are speed metrics.
- Write to /root/obs/slo-05-multi.promql the condition that the 1-hour and 5-minute burn rates each exceed 14.4, joined with and. Do not ignore a short window that has recovered.
- Write a multi-window alert rule named ErrorBudgetBurnFast in /etc/prometheus/rules/slo.yml. Include alert, expr, for, labels.severity, and annotations.summary, and check it with promtool check rules. Alert delivery and the feature deployment allow policy are separate matters.
- After reloading Prometheus, confirm through the rules API that the alert is registered. The fact that it is registered alone does not verify actual firing or operational response.
- Read /opt/lab/slo_release/contract.md and complete /root/obs/slo-gate.py. From one synthetic observation JSON on stdin, compute the period budget usage ratio and the burn rates of the two windows. Insufficient data is hold, a budget usage ratio of 1 or more is freeze, apart from that if both windows exceed 14.4 it is freeze, and the rest is allow. The output has five fields: decision, reason, budget_used, burn_hour, and burn_five_minutes, and the exit codes are allow=0, freeze=2, hold=3. The order of data checks and the detailed reasons follow the execution contract.
Notes
- Starter file: /opt/lab/slo_release/starter.py. If it does not exist, run mkdir -p /root/obs and then copy it to slo-gate.py.
- View input: python3 /opt/lab/slo_release/evaluate.py sample healthy
- Full verification: python3 /opt/lab/slo_release/evaluate.py check /root/obs/slo-gate.py
- Guide to the time-series checks for steps 05 and 06: /opt/lab/slo_release/promql.md. A false condition must return an empty vector, not an element with value 0, for the alert to turn off.
- Prometheus's roughly 12-hour synthetic history and the 30-day synthetic aggregate in the program input are different data. They are not real operational records.
- When there is no data, the number is null, not 0. A monthly ratio of 20% is 0.2, not 20.
- The program does not connect to external services or execute real deployments or rollbacks. allow means passing this synthetic policy, not a guarantee of operational safety.
Define the SLI as a query
Success ratio SLI → /root/obs/slo-01-sli.promql
This synthetic service counts only observed 5xx as failure. Divide the sum of the rate of successful requests by the sum of the total rate. Do not unconditionally classify 4xx of other services as success.
The 30-day error budget in minutes
Time-based 30-day budget (minutes) → /root/obs/slo-02-budget.txt
This is a time-based 99.9% SLO. Multiply the number of minutes in 30 days by the allowed downtime ratio and write it to one decimal place. Do not confuse it with a request-count budget.
Observe the 12-hour burn rate
Observe the 12-hour burn rate → /root/obs/slo-03-consumed.promql
Divide the 12-hour error rate by the allowed error ratio 0.001. Exceeding 1 means the speed in that window is faster than the allowed speed, not that the 30-day budget has already been used up.
Observe the 1-hour burn rate
Observe the 1-hour burn rate → /root/obs/slo-04-burn1h.promql
Change the window to 1h with the same error definition. Both the long window and the short window are burn rates, and you distinguish them from budget usage.
Multi-window condition
Multi-window condition → /root/obs/slo-05-multi.promql
Join with and the conditions that 1 hour and 5 minutes each exceed 14.4. You must remove false elements. Even a 0 from a bool comparison can turn the alert on. Use the time-series checks in /opt/lab/slo_release/promql.md to confirm recovery and missing data.
Write the alert rule
ErrorBudgetBurnFast multi-window alert rule → /etc/prometheus/rules/slo.yml
Include alert, expr, for, labels.severity, and annotations.summary. for: 0s evaluates the two-window condition with no additional wait for persistence. A long for can delay even the alert for a serious failure.
Confirm the rule is registered
Confirm that the ErrorBudgetBurnFast rule is registered in Prometheus
After the configuration check, reload and confirm the name in the rules API. Even if it is inactive, the registration check passes, but it does not replace a test of actual firing and alert delivery.
Verify the deployment decision as code
Decision program → /root/obs/slo-gate.py
Implement the data checks of the execution contract first. Leave a monthly budget overrun, a recent outage in the two windows, and unconfirmed data with different reasons and exit codes.