TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

The misleading denominator behind 100% success

Continue in TT Lab

One-line summary

The first question for an SLI is not the formula but what you count as one event and which failures you are missing. Dividing wrongly instrumented data exactly does not make it a correct user-experience metric.

Why this was needed

A hypothetical self-service platform provides a database issuance API. On success it returns 201, and on a dependency outage it returns 503. The owner incremented the counter only at the end of the success handler. Even if you send 2 healthy attempts and 2 failed attempts, only 2 successes and 2 totals remain in the metric. The dashboard's success rate is 100%, but the actual response record says 50%. At this point, raising the threshold from 99.9% to 99.99% cannot reveal the problem.

The lab that follows reproduces this incident. Instead of fake image captures, it runs a Python HTTP server inside a Pod and really receives 201, 503, and 504. The failure is chosen with a lab-only header called X-Lab-Scenario. It does not bring down a real external database. Only when you explain the layer that really runs and the cause that was virtually configured separately can you know the scope of the evidence.

How it works

Decide the unit of one event first

Google SRE's guide to implementing SLOs provides a starting point for designing measurable metrics and targets based on the actions users care about. Here we define one HTTP attempt to the issuance API as the unit. If a failed attempt a is followed by a successful retry b, that is two events. If you change it to "it was eventually issued, so one success," attempt-based metrics and user-task-based metrics get mixed. Neither is always right; they answer different questions.

The lab's scope.json is a hypothetical contract that treats only the 2xx and 5xx of /provision as valid attempts and excludes 4xx and the health check path. Do not copy this classification as is in production. For example, if a 429 occurs because of the service's insufficient capacity, excluding it as the user's fault can be unjust. You can prevent changing things so the metric merely looks better only if you agree on the validity of requests and the service's responsibility and also observe the exclusion ratio.

Separate good availability from good latency

In the lab, eligible is the attempts that go into the denominator, available is the 2xx among them, and fast is the responses among them that succeed and have a server processing time of 1000ms or less. 1000ms is included and 1000.1ms is excluded. A fast 503 is not fast. A slow 201 is available but not fast. The two SLIs use the same denominator but have different numerators.

Lab event eligible available fast
Healthy 201, 10ms true true true
Slow 201, 1100ms true true false
Failed 503, 10ms true false false
Input error 400 false false false

State the measurement location explicitly too. Server processing time does not include all of the client network round trip and connection wait. Do not stretch the fact that this lab measured server time into proof of the users' whole response time. If the real goal is latency as seen in a browser, you need observation at that location too.

Send failures directly and compare the two records

The Prometheus instrumentation principles are grounds for thinking about counters, failures, and label design together. The lab server does not put user IDs or request IDs in the metrics and exposes only four kinds: observed, eligible, available, and fast. Request IDs are handled in logs that track unique events, and metric labels are kept to a bounded set. This is to avoid a design in which per-user time series grow endlessly as users increase.

The HTTP replay matches seven responses with observed=7 and compares among them 4 in the denominator, 2 available successes, and 1 fast success. It does not count the /metrics scrape itself again as a business request. Rather than only checking the output that the metric is 100%, you connect the fact that you sent failures and the fact that the failures went into the denominator. It runs again with the current code, so submitting only an old healthy-result file is not enough.

What it looks like in the field

In another hypothetical incident, the collector sent the same log twice. If the attempt_id and the content are the same, it is a collection duplicate. Classify it only once and keep the duplicate count separately. If the ID is the same but the status code differs, you do not know which to trust, so you surface the conflict with a ValueError. If you arbitrarily pick the last value, a real failure can be covered by a success. On the other hand, a retry run with a different ID is a separate attempt, so you count both.

This rule is about small lab JSON events. If several services in production issue IDs, you also need the scope of collision prevention and a retention period. We do not claim that this task implemented exactly-once processing or permanent deduplication in a distributed environment.

What to read next

Next you read how much failure to allow with data whose denominator is verified, and what you cannot decide when the data is empty. In the lab that follows, you connect the classification function, real HTTP, duplicate handling, the error budget, and counterexamples you design yourself.