Why You Must Not Put user_id in a Label
In one line
The number of time series of a metric is the product of the label value combinations. If you choose one label wrongly, the time series become millions, and at that moment the monitoring system dies before the thing it observes.
Why this matters
You attach labels to http_requests_total.
http_requests_total{method, status, endpoint}
method: 5 values (GET, POST, PUT, DELETE, PATCH)status: 8 valuesendpoint: 40 values
The number of time series is 5 × 8 × 40 = 1,600. That is fine.
Then a request comes in saying "we want to see it per user", and you add user_id.
If there are 100,000 users,
5 × 8 × 40 × 100,000 = 160 million.
Prometheus keeps an index in memory for each time series. The server dies from OOM, and then the means of observing the outage disappears because of the outage.
Labels that cause a cardinality explosion
These are things you must never put in labels.
| Label | Why it is dangerous |
|---|---|
user_id, session_id |
Grows with the number of users. There is no upper bound |
request_id, trace_id |
Unique per request. Unbounded |
email, ip |
Effectively unbounded + personal data |
timestamp |
A new time series for every moment (the contradiction of putting time into a time series) |
url (including the query string) |
?page=1, ?page=2… unbounded |
| The full error message | Unbounded if values get mixed into the message |
Especially the last one. If you use a message with an IP and port in it, such as error="connection to 10.0.3.17:5432 timed out", as a label,
the combinations explode. Use a classified value such as error="db_timeout".
The criterion — how many values does this label have
Before adding a label, ask yourself.
- How many possible values are there? You must be able to answer by counting. "A lot" is not an answer.
- Does it grow over time? If it grows, is there an upper bound?
- Will you create real alerts from this label? If you think you will look at it once on a dashboard, it should go to logs or traces, not metrics.
As a rule of thumb, aim for 100 values or fewer per label and 10,000 time series or fewer per metric.
Always normalize URLs
# 나쁨 — 주문 개수만큼 시계열
endpoint="/api/orders/8f3a91"
# 좋음 — 라우트 패턴
endpoint="/api/orders/:id"
If you use the framework's route definitions, it is normalized automatically. If you write the string yourself, it will surely explode. This is not a matter of mistakes but a matter of time.
Use the three signals separately
There are times when you really need high-cardinality information. Questions like "why was this user's request slow?" That is not the job of metrics.
| Signal | Cardinality | The question it answers |
|---|---|---|
| Metrics | Must be low | "How much / how fast?" Trends and alerts |
| Logs | Can be high | "What happened then?" Individual events |
| Traces | Can be high | "Where did the time go?" The path of one request |
You receive an alert from metrics → pin down the time → dig into individual requests with logs and traces. The three signals are not substitutes but an investigation order.
If it has already exploded
- Find which metric it is. The TSDB status page in Prometheus shows a ranking of the metrics and labels with the most time series.
- Drop it at the collection stage. With relabel configuration, remove the problem label or drop the metric itself. You do not need to wait for an application deployment.
- Fix the application. The fundamental solution is not to attach the label.
- Adjust retention and sharding. This is for putting out the immediate fire, not a solution.
What it looks like in the field
- Prometheus periodically OOMs → suspect a newly added label.
- Dashboard queries take 30 seconds → too many time series, so the scan range is large.
- Memory rises in steps right after a deployment → that deployment added a label.
The order for reversing it after it has already blown up
You always learn about cardinality after an incident. Prometheus uses up all its memory and dies, or queries time out and never come back. There is an order for that moment.
First, count what takes up how much. Prometheus's status screen (/tsdb-status)
shows the metrics and labels that created the most time series. You can also see it with commands.
topk(10, count by (__name__)({__name__=~".+"}))
count(app_request_duration_seconds_bucket)
count(count by (user_id)(app_request_total))
The last line is the actual number of distinct values of that label. If this number is in the thousands, that one label is the cause.
Block it right before storage. Fixing and deploying the application is the right answer, but it takes time. In the meantime, remove the label in the scrape configuration or discard the metric entirely.
metric_relabel_configs:
- source_labels: [__name__]
regex: 'app_request_total'
target_label: user_id
replacement: '' # 라벨 값을 비운다
- source_labels: [__name__]
regex: 'debug_.*'
action: drop # 이 지표는 아예 저장하지 않는다
Time series that are already stored do not disappear on their own. Even if no new ones come in, they remain
for the retention period and use memory. If it is urgent, delete them with the admin API (--web.enable-admin-api must be
on, and deletion cannot be undone).
Histograms multiply silently. One bucket is one time series. If you multiply a histogram with 12 buckets by three labels, it becomes tens of thousands in an instant. If you cannot reduce the labels, reduce the number of buckets first.
What stops the same mistake is a pre-deployment check. Extract the metric names and label lists from the code, and block in CI if a label not on the allowlist is attached. If you rely on human attention, it will surely slip back in during a busy week.
What you will look at in the next check
In the following quiz, you calculate the number of time series as a product of label values, and distinguish the roles of metrics versus logs and traces. Also judge which to apply first in a situation where the explosion has already started: mitigation at the collection stage or a fundamental fix in the application.