One More Label and the Series Count Went Up a Hundredfold
Goal
You measure what the time series are being spent on now, predict in advance what the cost will be when you add a label, then actually start it up to confirm, pin the budget cap with a check script, and decide what to keep and what to discard.
Why it matters
Cardinality is a cost that nobody looks at until an incident happens. Adding one label is a single line of code, but if that label's values grow with the number of users, the time series grow as a product, and storage, memory, and query time all grow together. And that cost shows up in the form of "it got slow", so it is hard to find the cause. So you have to make two things a habit — multiplying out the number of combinations before adding a label, and pinning a per-metric time series cap with a check so that the pipeline, not a person, blocks it. The criterion for deciding what to discard is 'what do we decide with this label', and you must also check whether the question that the discarded axis used to answer can be answered by logs or traces instead.
Steps
- Create
/root/obs-cardinality-budget/inventory.tsv. For each metric thatjob="shop-api"exports, write the number of time series in two columns,<지표이름> <시계열 수>(the metric name, a tab, and the series count), listing all of them in descending order (five lines). The reason for narrowing to this one job rather than the whole Pod is that the metrics Prometheus itself emits keep growing even during the lab, which shifts the baseline. - Create
/root/obs-cardinality-budget/labels.tsv. It has three lines, and each line is<라벨이름> <서로 다른 값의 수>(the label name, a tab, and the number of distinct values). Countleandhandlerfrom the histogram metrichttp_request_duration_seconds_bucket, andstatusfrom the counterhttp_requests_total(the job is shop-api for both). Write them in descending order. Then write two lines in/root/obs-cardinality-budget/02-note.txt:top_label=<1위 라벨 이름>(the name of the top label) andproduct=<le 값 수 × handler 값 수>(the number of le values times the number of handler values). - Write three lines in
/root/obs-cardinality-budget/cost.txt.series=is the number of time series ofjob="shop-api",samples_per_day=is the number of samples those time series leave per day, andbytes_per_day=is the bytes those samples occupy. This Pod's scrape interval is 15 seconds, and one sample is taken as 2 bytes on average (the conservative end of the 1 to 2 bytes the Prometheus storage documentation mentions). The three values must fit together by multiplication. - You are about to export a new metric
checkout_requests_total. It has three labels:region(ap1, ap2, us1, eu1),tier(free, pro, team), andendpoint(cart, pay, refund, ship, track). Write two lines,formula=andpredicted_series=, in/root/obs-cardinality-budget/forecast.txt. formula is a multiplication expression (for example2*3*4), and predicted_series is its result. Do not start anything yet. - Create
/root/obs-cardinality-budget/exporter.py. If you give--printas an argument, it prints the exposition text to standard output and exits, and with no arguments it serves/metricsat 127.0.0.1:9102. It emits onecheckout_requests_totalline for each combination of the three labels from step 4. If the environment variableUSERSis greater than 0, theuser_idlabel is attached. Then add a scrape job namedcardinality-labto/etc/prometheus/prometheus.ymland have the configuration reloaded. After the first scrape finishes, write one linemeasured_series=<수>(the placeholder is the number) in/root/obs-cardinality-budget/measured.txt. - Start the exporter again with
USERS=100to attach theuser_idlabel. Write two lines in/root/obs-cardinality-budget/explode.tsv. Each line has three tab-separated columns,<이름> <시계열 수> <하루 바이트>(name, series count, bytes per day), where the name in the first line isbefore(without user_id) and the second line isafter(with user_id). Calculate bytes per day the same way as in step 3 (series × 5760 × 2). - Create
/root/obs-cardinality-budget/budget.sh. When called asbash budget.sh <상한표>(the argument is the cap table), for each line of the cap table (<지표이름> <최대 시계열 수>, the metric name and the maximum series count) it counts the current number of time series and prints one line per metric,<지표이름> <지금> <상한> OKor... OVER(metric name, current count, cap), and exits with code 1 if even one exceeds its cap and 0 otherwise. Then write the cap forcheckout_requests_totalas 500 in/root/obs-cardinality-budget/caps.tsv, run it once with that table, and save the output to/root/obs-cardinality-budget/gate-out.txt. - The time series cap is 500. Write four lines in
/root/obs-cardinality-budget/decision.txt. Inkeep=, write the label names to keep separated by commas (for exampleregion,tier); indrop=, the label names to drop separated by commas; inprojected_series=, the number of time series calculated from only the kept labels; and inlost_question=, the question that can no longer be answered once those labels are dropped, in at least 30 characters. The numbers of label values are region 4, tier 3, endpoint 5, and user_id 100.
Notes
- The working directory is
/root/obs-cardinality-budget. - TSDB status:
curl -s 'http://127.0.0.1:9090/api/v1/status/tsdb?limit=5' | jq— look atheadStats,seriesCountByMetricName, andlabelValueCountByLabelName. - Apply configuration:
curl -X POST http://127.0.0.1:9090/-/reload. The scrape interval is 15 seconds, so it takes a little while to get the first value. - Common mistake: not treating an empty
count()result as 0, so the gate ends with an error. - Common mistake: multiplying on the assumption that the label values are independent when in fact only some combinations exist — the forecast is an upper bound, and if the measured value is smaller than that, the difference is also information.
- Storage · TSDB status API · Instrumentation practices · Naming
Count what one service's time series are spent on
Create /root/obs-cardinality-budget/inventory.tsv. For each metric that job="shop-api" exports, write the number of time series in two columns, <지표이름> <시계열 수> (the metric name, a tab, and the series count), listing all of them in descending order (five lines). The reason for narrowing to this one job rather than the whole Pod is that the metrics Prometheus itself emits keep growing even during the lab, which shifts the baseline.
If you throw count by (__name__)(last_over_time({job="shop-api"}[1m])), the number of time series per metric comes out. You will immediately see that one histogram creates as many time series as it has buckets.
Which label creates the most values
Create /root/obs-cardinality-budget/labels.tsv. It has three lines, and each line is <라벨이름> <서로 다른 값의 수> (the label name, a tab, and the number of distinct values). Count le and handler from the histogram metric http_request_duration_seconds_bucket, and status from the counter http_requests_total (the job is shop-api for both). Write them in descending order. Then write two lines in /root/obs-cardinality-budget/02-note.txt: top_label=<1위 라벨 이름> (the name of the top label) and product=<le 값 수 × handler 값 수> (the number of le values times the number of handler values).
You count how many values a label has with count(count by (<라벨>) (...)) (the placeholder is the label name). To look at only the time series alive now, wrap it in last_over_time(<선택자>[1m]) (the placeholder is the selector) — otherwise you also count the backfilled past time series and the value wobbles. The final product must equal the number of time series of the histogram you saw in step 1.
How much do this service's metrics consume per day
Write three lines in /root/obs-cardinality-budget/cost.txt. series= is the number of time series of job="shop-api", samples_per_day= is the number of samples those time series leave per day, and bytes_per_day= is the bytes those samples occupy. This Pod's scrape interval is 15 seconds, and one sample is taken as 2 bytes on average (the conservative end of the 1 to 2 bytes the Prometheus storage documentation mentions). The three values must fit together by multiplication.
The number of time series is count(last_over_time({job="shop-api"}[1m])). A day is 86400 seconds, so with a 15-second interval the number of samples per time series is fixed. This number is the unit price of 'the cost that grows when you add one label'.
Predict first, before measuring
You are about to export a new metric checkout_requests_total. It has three labels: region (ap1, ap2, us1, eu1), tier (free, pro, team), and endpoint (cart, pay, refund, ship, track). Write two lines, formula= and predicted_series=, in /root/obs-cardinality-budget/forecast.txt. formula is a multiplication expression (for example 2*3*4), and predicted_series is its result. Do not start anything yet.
The number of label combinations is the number of time series. If the values are independent of each other, multiply the number of values of each label. The reason to write the forecast first is that if you adjust the number after measuring, you cannot tell what you had thought wrongly.
Start the exporter and confirm the forecast
Create /root/obs-cardinality-budget/exporter.py. If you give --print as an argument, it prints the exposition text to standard output and exits, and with no arguments it serves /metrics at 127.0.0.1:9102. It emits one checkout_requests_total line for each combination of the three labels from step 4. If the environment variable USERS is greater than 0, the user_id label is attached. Then add a scrape job named cardinality-lab to /etc/prometheus/prometheus.yml and have the configuration reloaded. After the first scrape finishes, write one line measured_series=<수> (the placeholder is the number) in /root/obs-cardinality-budget/measured.txt.
For the server, the ThreadingHTTPServer of http.server is enough. The exposition format is 이름{라벨="값",...} 값 on each line (metric name, labels in braces, then the value). Apply the configuration with curl -X POST http://127.0.0.1:9090/-/reload, and since the scrape interval is 15 seconds, you have to wait a little. Count the number of time series with count(checkout_requests_total).
One label multiplies the time series a hundredfold
Start the exporter again with USERS=100 to attach the user_id label. Write two lines in /root/obs-cardinality-budget/explode.tsv. Each line has three tab-separated columns, <이름> <시계열 수> <하루 바이트> (name, series count, bytes per day), where the name in the first line is before (without user_id) and the second line is after (with user_id). Calculate bytes per day the same way as in step 3 (series × 5760 × 2).
The number of time series comes straight out if you run the exporter with --print and count the lines that start with checkout_requests_total — the server does not need to be up. The difference between the two lines is 'the values of one label'.
Pin the cap down in code
Create /root/obs-cardinality-budget/budget.sh. When called as bash budget.sh <상한표> (the argument is the cap table), for each line of the cap table (<지표이름> <최대 시계열 수>, the metric name and the maximum series count) it counts the current number of time series and prints one line per metric, <지표이름> <지금> <상한> OK or ... OVER (metric name, current count, cap), and exits with code 1 if even one exceeds its cap and 0 otherwise. Then write the cap for checkout_requests_total as 500 in /root/obs-cardinality-budget/caps.tsv, run it once with that table, and save the output to /root/obs-cardinality-budget/gate-out.txt.
Count the number of time series with count(<지표>) (the placeholder is the metric). If there is no result, you must treat it as 0 — you can give a default with jq's //. The exit code must be issued only once at the end, so collect it in a variable. The grader runs this script with two cap tables it makes itself and confirms both the pass and the fail.
What to discard to fit within the budget
The time series cap is 500. Write four lines in /root/obs-cardinality-budget/decision.txt. In keep=, write the label names to keep separated by commas (for example region,tier); in drop=, the label names to drop separated by commas; in projected_series=, the number of time series calculated from only the kept labels; and in lost_question=, the question that can no longer be answered once those labels are dropped, in at least 30 characters. The numbers of label values are region 4, tier 3, endpoint 5, and user_id 100.
The criterion for choosing which label to drop is 'what do we decide with this label'. A label that does not change a decision is expensive. Per-person investigation is a question that logs and traces can answer, so even if you drop that axis from metrics, you do not lose the ability to investigate entirely.