TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

The dashboard said p90 was 0.48s; the logs said 0.42s

Continue in TT Lab

Goal

You aggregate the same requests with two versions that have different bucket boundaries to measure the estimation error of histogram_quantile yourself, check what numbers come from an expression that drops le, an expression that mixes versions, and an expression that leaves out rate, and then design buckets that fit an SLO.

Why it matters

A histogram remembers samples only as bucket counts. So a quantile is the result of linear interpolation that assumes the samples are spread evenly within a bucket, and the size of the error is set by the width of the bucket in which the rank falls. A tail that goes beyond the boundary is not visible at all. You can trust an SLO only if you can tell whether to suspect the query or the bucket design when a dashboard number differs from the logs.

Prepared environment

python3 /opt/fixtures/pca_histogram_lab.py init puts the v2 Pod's request log (access-log-v2.csv) and a README in /root/pca-histogram/. The time series are http_request_duration_seconds_bucket/_count/_sum, and the labels are route (/checkout, /search), version (v1, v2), and le. v1 has boundaries 0.1 0.5 1, and v2 has 0.05 0.1 0.2 0.3 0.4 0.5 1 2.5. It was scraped every minute from minute 0 to minute 60, and the current time is 60m.

python3 /opt/fixtures/pca_histogram_lab.py eval -f 파일 --at 60m (the placeholder is the file) evaluates an expression with the promtool 3.0.1 PromQL engine from the lab-k8s image. It does not start a Prometheus server, and the grader also recomputes your expressions and the reference expressions with the same engine and compares them against the numbers you wrote.

Steps

  1. From /root/pca-histogram/access-log-v2.csv, pick the 100 requests for /checkout from minute 56 to minute 60, sort them, and find p90 and p99 as the ceil(q×N)-th value (nearest-rank). Write them in /root/pca-histogram/01-raw.txt as two lines, raw_p90= and raw_p99= (in seconds). If the material is missing, run python3 /opt/fixtures/pca_histogram_lab.py init first.
  2. In /root/pca-histogram/02-v1-p90.promql, write the p90 expression over the last 5 minutes for version="v1", route="/checkout" (histogram_quantile(0.9, ...), rate(...[5m]), keep le). Look at the result with python3 /opt/fixtures/pca_histogram_lab.py eval -f /root/pca-histogram/02-v1-p90.promql --at 60m and write v1_p90= and error= (v1_p90 minus raw_p90) in /root/pca-histogram/02-v1.txt.
  3. In /root/pca-histogram/03-v2-p90.promql, write the same expression with version="v2" and check the result. In /root/pca-histogram/03-v2.txt, write v2_p90=, error= (v2_p90 minus raw_p90), and closer= (the version, v1 or v2, closer to the log's p90).
  4. Check the last-5-minutes p99 of /checkout for both v1 and v2 with eval, and in /root/pca-histogram/04-slow.promql, write the number of v1 /checkout requests that exceeded 1 second in the last 5 minutes as an expression that subtracts the le="1" bucket from increase(..._bucket{le="+Inf"}[5m]) (the le values differ, so a matching condition is needed). In /root/pca-histogram/04-p99.txt, write v1_p99=, v2_p99=, and slow_requests=.
  5. In /root/pca-histogram/05-broken.promql, write the v2 p90 expression that drops le with sum by (route), and in /root/pca-histogram/05-fixed.promql, write the v2 p90 expression that keeps both le and route. Eval both expressions and write broken_series= (the number of result time series), fixed_series=, and search_p90= (the value for /search) in /root/pca-histogram/05-le.txt.
  6. During a canary deployment, the dashboard combines the /checkout buckets without distinguishing version. In /root/pca-histogram/06-mixed.promql, write a p90 expression that selects only route="/checkout" without choosing a version, and count the kinds of le values with count by (le) (http_request_duration_seconds_bucket{route="/checkout"}). In /root/pca-histogram/06-mixed.txt, write mixed_p90= and distinct_le=.
  7. In /root/pca-histogram/07-lifetime.promql, write a p90 expression that sums the v2 /checkout cumulative buckets as they are, without rate. In /root/pca-histogram/07-lifetime.txt, write lifetime_p90= and window_p90= (the last-5-minutes v2 p90 from step 03), and compare which of the two shows more clearly the fact that it got slower at minute 31.
  8. In /root/pca-histogram/08-buckets.txt, write the new bucket boundaries on one line (ascending, at most 10 finite boundaries, without +Inf). There are three conditions — it includes the SLO boundary 0.3, the largest finite boundary is at least the maximum latency of 1.8 seconds, and the error against the log p90 is at most 0.01. Look at the result of re-aggregating the same requests with the new boundaries using python3 /opt/fixtures/pca_histogram_lab.py design, and write designed_p90= and under_300ms_ratio= in /root/pca-histogram/08-design.txt.

Notes

Counting the real p90 and p99 from the log

From /root/pca-histogram/access-log-v2.csv, pick the 100 requests for /checkout from minute 56 to minute 60, sort them, and find p90 and p99 as the ceil(q×N)-th value (nearest-rank). Write them in /root/pca-histogram/01-raw.txt as two lines, raw_p90= and raw_p99= (in seconds). If the material is missing, run python3 /opt/fixtures/pca_histogram_lab.py init first.

Tools that do linear interpolation (such as the numpy default) return a value between two samples. Here you pick the one actual sample that corresponds to the rank.

The p90 that the coarse buckets (v1) report

In /root/pca-histogram/02-v1-p90.promql, write the p90 expression over the last 5 minutes for version="v1", route="/checkout" (histogram_quantile(0.9, ...), rate(...[5m]), keep le). Look at the result with python3 /opt/fixtures/pca_histogram_lab.py eval -f /root/pca-histogram/02-v1-p90.promql --at 60m and write v1_p90= and error= (v1_p90 minus raw_p90) in /root/pca-histogram/02-v1.txt.

A quantile is linearly interpolated within a bucket. v1 has no boundary between 0.1 and 0.5, so it assumes that entire range is evenly spread.

Comparing with the dense buckets (v2)

In /root/pca-histogram/03-v2-p90.promql, write the same expression with version="v2" and check the result. In /root/pca-histogram/03-v2.txt, write v2_p90=, error= (v2_p90 minus raw_p90), and closer= (the version, v1 or v2, closer to the log's p90).

v2 has boundaries at 0.4 and 0.5. The narrower the bucket in which the rank falls, the lower the upper bound of the interpolation error.

A p99 trapped at the last finite boundary

Check the last-5-minutes p99 of /checkout for both v1 and v2 with eval, and in /root/pca-histogram/04-slow.promql, write the number of v1 /checkout requests that exceeded 1 second in the last 5 minutes as an expression that subtracts the le="1" bucket from increase(..._bucket{le="+Inf"}[5m]) (the le values differ, so a matching condition is needed). In /root/pca-histogram/04-p99.txt, write v1_p99=, v2_p99=, and slow_requests=.

If the rank falls in the +Inf bucket, histogram_quantile returns the largest finite boundary. The two bucket time series differ only in the le label, so ignoring(le) is needed.

Fixing a dashboard that dropped le

In /root/pca-histogram/05-broken.promql, write the v2 p90 expression that drops le with sum by (route), and in /root/pca-histogram/05-fixed.promql, write the v2 p90 expression that keeps both le and route. Eval both expressions and write broken_series= (the number of result time series), fixed_series=, and search_p90= (the value for /search) in /root/pca-histogram/05-le.txt.

histogram_quantile reads as buckets only the time series that have an le label. An empty result with no error is the most dangerous form.

When you combine two versions with different boundaries

During a canary deployment, the dashboard combines the /checkout buckets without distinguishing version. In /root/pca-histogram/06-mixed.promql, write a p90 expression that selects only route="/checkout" without choosing a version, and count the kinds of le values with count by (le) (http_request_duration_seconds_bucket{route="/checkout"}). In /root/pca-histogram/06-mixed.txt, write mixed_p90= and distinct_le=.

The le of the combined result is the union of the two versions' boundaries. At a boundary that exists on only one side, the cumulative counts get distorted and the engine forces monotonicity.

When you read cumulative buckets without rate

In /root/pca-histogram/07-lifetime.promql, write a p90 expression that sums the v2 /checkout cumulative buckets as they are, without rate. In /root/pca-histogram/07-lifetime.txt, write lifetime_p90= and window_p90= (the last-5-minutes v2 p90 from step 03), and compare which of the two shows more clearly the fact that it got slower at minute 31.

A counter's cumulative value holds all the requests since the process started. If the 30 minutes before it got slow are mixed in, the recent change is diluted.

Redesigning the buckets to fit the SLO

In /root/pca-histogram/08-buckets.txt, write the new bucket boundaries on one line (ascending, at most 10 finite boundaries, without +Inf). There are three conditions — it includes the SLO boundary 0.3, the largest finite boundary is at least the maximum latency of 1.8 seconds, and the error against the log p90 is at most 0.01. Look at the result of re-aggregating the same requests with the new boundaries using python3 /opt/fixtures/pca_histogram_lab.py design, and write designed_p90= and under_300ms_ratio= in /root/pca-histogram/08-design.txt.

The error comes from the width of the bucket in which the rank falls. You do not need to make everywhere dense; you only need to narrow the range where p90 sits and the SLO boundary.