TT Lab
Get started
Learn Learning paths Courses

Observability

Same 12 Hours, Availability of 98.9% and 85.0%

Continue in TT Lab

Goal

You measure the same 12 hours of data yourself with four SLI definitions to see how far the numbers diverge, pick one and declare it together with an objective, and then decide whether to allow or block a deployment using that definition.

Why it matters

What percent to set an SLO at is a question that comes after what to count has been decided. Depending on how you define good events and valid events, the same data gives 98.9% or 85.0%. If you put health checks in the denominator, the outage users experienced is diluted, and if you count a slow success as a success, the time users left does not remain in the record. If you change the unit of an event from requests to time, you end up counting a dawn outage and a daytime outage with the same weight. These choices come back later as whether a deployment is blocked or not, so whenever you change a definition, you need the habit of re-measuring the past period and confirming that you can live with that number.

Steps

  1. Write three lines in /root/obs-sli-events/01-events.txt. After metric=, write the name of the counter metric to use for counting events; after failure_label=, the name of the label in that metric that separates failure from success; and after valid_12h=, the number of valid events in the last 12 hours as an integer (the job is shop-api). Get the values by throwing queries yourself.
  2. In /root/obs-sli-events/a-request.promql, write a PromQL of one or more lines that gives 'the ratio of requests that are not 5xx over the last 12 hours', and write its result to /root/obs-sli-events/a-request.txt as a single number to four decimal places. You may leave comments (#).
  3. Measure the same 12 hours again with '2xx responses only are good events'. Write the query to /root/obs-sli-events/b-strict.promql and the value to /root/obs-sli-events/b-strict.txt to four decimal places. The denominator must be the same as in the previous step.
  4. Measure the 12 hours with the definition that treats 'requests that finished within 100 milliseconds' as good events. Write the query to /root/obs-sli-events/c-latency.promql and the value to /root/obs-sli-events/c-latency.txt to four decimal places. Then write in /root/obs-sli-events/04-note.txt, on one line, the fact that 'with this definition, you cannot look at status code and latency together in one query' and the reason — it must start with reason= and be at least 40 characters.
  5. Measure the same 12 hours with the definition that takes a 1-minute window as an event. If the 5xx ratio within a given minute is less than 1%, treat that minute as a good event. Write the query to /root/obs-sli-events/d-window.promql and the value to /root/obs-sli-events/d-window.txt to four decimal places. Then write in /root/obs-sli-events/05-badminutes.txt how many of the 12 hours (720 minutes) were bad minutes, as an integer.
  6. Create /root/obs-sli-events/budget.tsv. It has four lines with no header, and each line has three tab-separated columns, <id> <가용성> <30일 허용 다운타임 분> (the id, the availability, and the 30-day allowed downtime in minutes). The ids are a, b, c, and d in order, the availability is the value you wrote in the previous steps, and in the third column write, to one decimal place, how many minutes out of 30 days (43200 minutes) may be bad at that availability.
  7. Write four lines in /root/obs-sli-events/choice.txt. After sli=, the id you chose (one of a, b, c, d); after query=, the query file name of that definition (for example d-window.promql); after target=, the objective you will set as a decimal between 0 and 1; and after reason=, why that definition, in at least 60 characters. The grader actually throws the file that query= points to and checks that it matches sli=.
  8. Write two lines in /root/obs-sli-events/decision.txt. Each line is target=<목표> consumed=<소진비율> release=<allow|hold> (objective, consumption ratio, and the release decision), and the objective is 0.99 in the first line and 0.95 in the second. The consumption ratio is (1 − the measured availability) ÷ (1 − the objective), to three decimal places, and the decision is hold if the consumption ratio is 1 or more and allow if it is less. For the measured availability, use the value of the definition you chose in step 7.

Notes

Decide first what to count as one

Write three lines in /root/obs-sli-events/01-events.txt. After metric=, write the name of the counter metric to use for counting events; after failure_label=, the name of the label in that metric that separates failure from success; and after valid_12h=, the number of valid events in the last 12 hours as an integer (the job is shop-api). Get the values by throwing queries yourself.

You can see which metrics exist with promq "{__name__=~\"http.*\"}" or curl -s http://127.0.0.1:9090/api/v1/label/__name__/values | jq. Get the amount a counter grew over 12 hours with increase. For the label name, write only the name, not the values.

Request basis — count only 5xx as failure

In /root/obs-sli-events/a-request.promql, write a PromQL of one or more lines that gives 'the ratio of requests that are not 5xx over the last 12 hours', and write its result to /root/obs-sli-events/a-request.txt as a single number to four decimal places. You may leave comments (#).

A counter is cumulative, so you must convert it into an increase over the range. Aggregate the numerator and the denominator each with sum and then divide. To select 5xx by label matching, use a regular expression matcher.

How much does it change if 400-range is also counted as failure

Measure the same 12 hours again with '2xx responses only are good events'. Write the query to /root/obs-sli-events/b-strict.promql and the value to /root/obs-sli-events/b-strict.txt to four decimal places. The denominator must be the same as in the previous step.

400-range responses are usually the client's fault, so they tend not to be counted as service failures, but there are also 4xx that are our responsibility, as in an incident where authentication failures surge. See which way the number moves when you change the definition.

Is a slow success a success

Measure the 12 hours with the definition that treats 'requests that finished within 100 milliseconds' as good events. Write the query to /root/obs-sli-events/c-latency.promql and the value to /root/obs-sli-events/c-latency.txt to four decimal places. Then write in /root/obs-sli-events/04-note.txt, on one line, the fact that 'with this definition, you cannot look at status code and latency together in one query' and the reason — it must start with reason= and be at least 40 characters.

Histogram buckets are cumulative. Look at which bucket, with what le value, is 'the number of requests that finished at or below that value'. The total number of requests to use as the denominator is in the count series of the histogram. The key point is that the label sets of the two metrics are different.

Count minutes instead of requests

Measure the same 12 hours with the definition that takes a 1-minute window as an event. If the 5xx ratio within a given minute is less than 1%, treat that minute as a good event. Write the query to /root/obs-sli-events/d-window.promql and the value to /root/obs-sli-events/d-window.txt to four decimal places. Then write in /root/obs-sli-events/05-badminutes.txt how many of the 12 hours (720 minutes) were bad minutes, as an integer.

Make values at 1-minute intervals with the subquery [12h:1m], and if you put bool after the comparison operator, you get a series where true is 1 and false is 0. The average of that series is the ratio of good minutes. Count the bad minutes backward from 720.

Convert the four definitions into 30-day allowed downtime

Create /root/obs-sli-events/budget.tsv. It has four lines with no header, and each line has three tab-separated columns, <id> <가용성> <30일 허용 다운타임 분> (the id, the availability, and the 30-day allowed downtime in minutes). The ids are a, b, c, and d in order, the availability is the value you wrote in the previous steps, and in the third column write, to one decimal place, how many minutes out of 30 days (43200 minutes) may be bad at that availability.

The allowed downtime is (1 − availability) × 43200. How far the four numbers diverge is the heart of this step — changing one definition changes the time the operations team has to bear by several times.

Declare the chosen definition in a machine-readable way

Write four lines in /root/obs-sli-events/choice.txt. After sli=, the id you chose (one of a, b, c, d); after query=, the query file name of that definition (for example d-window.promql); after target=, the objective you will set as a decimal between 0 and 1; and after reason=, why that definition, in at least 60 characters. The grader actually throws the file that query= points to and checks that it matches sli=.

There is not just one right answer. However, the grounds for your choice must connect to the numbers from the previous steps — for example, if you choose the latency basis, you should choose only after confirming in the step 6 table whether the 30-day allowed downtime is bearable.

The same data, a different objective — allow the deployment?

Write two lines in /root/obs-sli-events/decision.txt. Each line is target=<목표> consumed=<소진비율> release=<allow|hold> (objective, consumption ratio, and the release decision), and the objective is 0.99 in the first line and 0.95 in the second. The consumption ratio is (1 − the measured availability) ÷ (1 − the objective), to three decimal places, and the decision is hold if the consumption ratio is 1 or more and allow if it is less. For the measured availability, use the value of the definition you chose in step 7.

A consumption ratio of 1 means 'the error budget available in this period has been used up exactly'. Even with the same data, the deployment decision flips depending on where you set the objective — which is why an objective is an agreement, not a number.