TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Does removing a label really aggregate values?

Continue in TT Lab

In one line

Dropping a label is not summing, and up=1 is not proof that all the needed information was preserved. You have to compare the meaning of the raw data with the query results.

Why this was needed

With too many metrics, it seems that erasing just the user label would solve it. If twelve counters look like one, the storage cost has gone down and the work seems done. But in an experiment where the total should be 78, the result became 1. The target status is still up and the recent error string is empty. Making the screen green and preserving the question the operator asks are different goals. This difference is easy to miss before going through a real collector.

How it works

metric_relabel_configs changes the labels of collected samples or excludes some samples. labeldrop is the action that removes the specified label name. If you erase that label from samples that different user_id values used to distinguish, the addresses with the same name and the same remaining labels collide. This operation is not an aggregation that adds up several counter values. The official documentation also says to take care that the uniqueness of time series is maintained after removing a label.

In this fixed input, the per-user values are 1 through 12 and the total is 78. In the real Prometheus 3.14.0 probe, after erasing user_id, the current number of time series was 1 and the query total was also 1. up was 1 and lastError was empty. We do not generalize this result into an API contract that "on a collision the first value is always preserved." The key point of this observation, with the version, input, and time fixed, is that labeldrop did not sum up the 78. With other inputs or versions, a different error or rejection form may appear, so you have to check the raw response and the actual result directly.

The first verdict checker of this experiment failed by assuming that a collision would necessarily give up=0. When the actual result differed from expectation, we did not force the server to down. We left the code at the time and a read-only observation, and on a new VM compared the input and result again. Good verification is not making the state you expected but explaining what actually happened. When a test is wrong, the test's assumptions are also subject to correction.

The next fix is to drop the entire problem request metric. The raw response is still 13 samples, but after collection only one business gauge remains and fits within the limit. up is normal again. But the request counter query becomes an empty vector. You must not report this as zero requests. Having no information to observe and having a sample whose value is 0 are different. From the query alone it can be hard to tell whether the filter is an intended exclusion or a gap caused by an outage.

Change What looks good on screen The loss you must check
Raising sample_limit Collection succeeds again The high number of label combinations remains
Dropping the distinguishing label The current number of time series goes down Summation of distinct values and preservation of meaning are not guaranteed
Excluding the whole metric family It is within the limit and up is normal The metric that answers the needed business question may vanish
Designing the dimensions at the source and aggregating The needed total is exposed with few time series Questions about the dropped detail dimension can no longer be answered

At the end, you change the exporter's instrumentation model. Instead of the detailed per-synthetic-user values, it exposes a counter aggregated at the route level, 78, and does not emit the user_id label from the start. With one request counter and one business gauge, the raw response is two samples from the start. Even if you remove the temporary filter and put sample_limit back to 8, collection succeeds and the request total stays at 78. This is the step that distinguishes quietly discarding behind a filter from designing the needed dimension at the source.

What it looks like in the field

A dashboard's sum adds up the query results. Saving that query does not mean the per-user time series already collected are erased from the TSDB. In the lab you look side by side at the current number of time series and the query at a past evaluation time. After the improvement there is currently one time series, but if you specify the time when you raised the limit and collected 12, you can check again the 12 and the total 78 at that time. We do not conclude from the label list of the series API alone that actual samples existed at that time, and we leave together the real PromQL result with the evaluation time specified. A metadata list and time series samples are not the same evidence.

This lab does not turn on the TSDB delete API or delete data. This is so as not to claim that fixing the source instrumentation automatically erases past sensitive information. When dealing with a retention policy or deletion in production, you have to review separate permissions, backups, auditing, and the scope of impact. Here, every user ID is a synthetic value, and we concentrate on safely separating two different problems: the current instrumentation fix and the preservation of past data.

Whether current information is properly preserved is also not settled with a single number. You read together the exporter's raw response, the currently applied configuration, the most recent scrape time, the current counter total, and an independent control group. You check that the services' run identifiers are kept, recording that state was not reset by a restart. Being normal in a short sample and having verified long-term performance and loss rate are different, and this small experiment does not stand in for a large-scale load test.

What you will do in the next lab

You compare in turn raising the limit, dropping the colliding label, excluding the metric, and aggregating at the source. In the final report you write why up alone cannot guarantee information preservation, the business question you kept, the detail dimension you intentionally dropped, and the point that the current fix does not erase the past. The habit of verifying "collection success" and "an observation whose meaning is right" separately is the product of this lab.

Official documentation