TT Lab
Get started
Learn Learning paths Courses

Observability

Why You Must Not Put user_id in a Label

Continue in TT Lab

In one line

The number of time series of a metric is the product of the label value combinations. If you choose one label wrongly, the time series become millions, and at that moment the monitoring system dies before the thing it observes.

Why this matters

You attach labels to http_requests_total.

http_requests_total{method, status, endpoint}

The number of time series is 5 × 8 × 40 = 1,600. That is fine.

Then a request comes in saying "we want to see it per user", and you add user_id. If there are 100,000 users,

5 × 8 × 40 × 100,000 = 160 million.

Prometheus keeps an index in memory for each time series. The server dies from OOM, and then the means of observing the outage disappears because of the outage.

Labels that cause a cardinality explosion

These are things you must never put in labels.

Label Why it is dangerous
user_id, session_id Grows with the number of users. There is no upper bound
request_id, trace_id Unique per request. Unbounded
email, ip Effectively unbounded + personal data
timestamp A new time series for every moment (the contradiction of putting time into a time series)
url (including the query string) ?page=1, ?page=2… unbounded
The full error message Unbounded if values get mixed into the message

Especially the last one. If you use a message with an IP and port in it, such as error="connection to 10.0.3.17:5432 timed out", as a label, the combinations explode. Use a classified value such as error="db_timeout".

The criterion — how many values does this label have

Before adding a label, ask yourself.

  1. How many possible values are there? You must be able to answer by counting. "A lot" is not an answer.
  2. Does it grow over time? If it grows, is there an upper bound?
  3. Will you create real alerts from this label? If you think you will look at it once on a dashboard, it should go to logs or traces, not metrics.

As a rule of thumb, aim for 100 values or fewer per label and 10,000 time series or fewer per metric.

Always normalize URLs

# 나쁨 — 주문 개수만큼 시계열
endpoint="/api/orders/8f3a91"

# 좋음 — 라우트 패턴
endpoint="/api/orders/:id"

If you use the framework's route definitions, it is normalized automatically. If you write the string yourself, it will surely explode. This is not a matter of mistakes but a matter of time.

Use the three signals separately

There are times when you really need high-cardinality information. Questions like "why was this user's request slow?" That is not the job of metrics.

Signal Cardinality The question it answers
Metrics Must be low "How much / how fast?" Trends and alerts
Logs Can be high "What happened then?" Individual events
Traces Can be high "Where did the time go?" The path of one request

You receive an alert from metrics → pin down the time → dig into individual requests with logs and traces. The three signals are not substitutes but an investigation order.

If it has already exploded

What it looks like in the field

The order for reversing it after it has already blown up

You always learn about cardinality after an incident. Prometheus uses up all its memory and dies, or queries time out and never come back. There is an order for that moment.

First, count what takes up how much. Prometheus's status screen (/tsdb-status) shows the metrics and labels that created the most time series. You can also see it with commands.

topk(10, count by (__name__)({__name__=~".+"}))
count(app_request_duration_seconds_bucket)
count(count by (user_id)(app_request_total))

The last line is the actual number of distinct values of that label. If this number is in the thousands, that one label is the cause.

Block it right before storage. Fixing and deploying the application is the right answer, but it takes time. In the meantime, remove the label in the scrape configuration or discard the metric entirely.

metric_relabel_configs:
  - source_labels: [__name__]
    regex: 'app_request_total'
    target_label: user_id
    replacement: ''            # 라벨 값을 비운다
  - source_labels: [__name__]
    regex: 'debug_.*'
    action: drop               # 이 지표는 아예 저장하지 않는다

Time series that are already stored do not disappear on their own. Even if no new ones come in, they remain for the retention period and use memory. If it is urgent, delete them with the admin API (--web.enable-admin-api must be on, and deletion cannot be undone).

Histograms multiply silently. One bucket is one time series. If you multiply a histogram with 12 buckets by three labels, it becomes tens of thousands in an instant. If you cannot reduce the labels, reduce the number of buckets first.

What stops the same mistake is a pre-deployment check. Extract the metric names and label lists from the code, and block in CI if a label not on the allowlist is attached. If you rely on human attention, it will surely slip back in during a busy week.

What you will look at in the next check

In the following quiz, you calculate the number of time series as a product of label values, and distinguish the roles of metrics versus logs and traces. Also judge which to apply first in a situation where the explosion has already started: mitigation at the collection stage or a fundamental fix in the application.