TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

The business is healthy, but scraping fails

Continue in TT Lab

In one line

A label is not a note that adds explanation but the address of a time series. A sample limit protects the collection budget, but it does not judge what information you lost.

Why this was needed

While improving a payment service, a user identifier was attached to the request metrics. It seemed it would be convenient for finding a particular customer's problem. It looked fine in a small development environment, but as users grew, the monitoring target was shown as down. The service's health check address returns HTTP 200 and the payment process is alive. If you think it is a network failure and restart, it seems to get better for a while, but when the same label combinations accumulate again, the problem comes back. First you have to read at which stage the collection was rejected.

In Prometheus, one time series is distinguished by the metric name and the full label set. Even for a counter of the same name, if user_id is different, it is a different time series. It is not summed automatically just because route is the same. A change a person thinks of as adding one label column becomes an increase in the number of addresses in storage. Not only the number of users but also the actual combinations with route, status code, and instance matter. The product of possible combinations is a starting point for estimating the upper bound, not always the exact number observed.

How it works

This lab uses at most twelve synthetic users. It collects no real personal information and does not create excessive memory load. The baseline exporter exposes request counters for four users and one business status gauge. One response has five samples. If you increase the users to twelve, even with the same metric names it becomes twelve counter samples and one gauge, thirteen samples in total. A separate fixed control target keeps emitting the same values to the end.

Observation point The question it answers in this experiment
The exporter's /healthz and /metrics Is the process responding normally, and what raw data does it provide?
The scrape configuration currently applied What targets, limits, and filters are actually in use?
The target API's health and lastError Did the recent collection succeed, and if it failed, why?
The current PromQL result With what labels and values are the stored samples being read?

sample_limit limits the number of samples accepted in a single scrape. It is a setting where, if the number after metric relabeling exceeds the limit, the whole scrape fails. It is not a page size that stores only the first eight and trims the rest normally. If you set the lab's limit to 8, the five-sample baseline passes but the thirteen-sample response is rejected. Connect the recent error that the limit was exceeded with up=0, but also keep the business HTTP 200 and the control target's up=1.

If you raise sample_limit to 16, this small response comes in again. This is a useful experiment for narrowing down the cause, but it is not evidence that the instrumentation design got better. The request time series are still twelve. In a production environment where users keep growing, if you only raise the limit, you can only postpone the point at which the next limit is reached. 8 and 16 are lab numbers for observing the principle and are not production-recommended values to apply to every service. The real budget must be measured together with the scrape interval, number of targets, kinds of metrics, retention period, and query patterns.

Metric names also convey the unit of the signal. The request count is a cumulative counter and its name is pca_checkout_requests_total. The current business health is the gauge pca_business_ok. The total of the cumulative counter values is a value for checking information preservation in this experiment, not throughput per second. To get throughput per second, you have to solve the separate problem of first computing a rate that accounts for each counter's reset and then aggregating. Not mixing the two concepts with a small static input is the intent of this design.

What it looks like in the field

Labels whose range of values grows, such as user ID, order number, and full URL, look like a convenient search feature. But metrics have a different role from a store that keeps the details of every event. Design first the dimensions needed for the operational questions, like a path template or a bounded status classification, and consider linking the context of individual events to appropriate logs and traces. Once sensitive identifiers enter metrics, the scope of access control and of retention and deletion also becomes complicated. A synthetic value like u1 here does not mean it is fine to put in real personal information.

In an incident report, do not write only the word down. If you write "business HTTP is normal, control collection is normal, the raw response of that job is 13 samples, and the scrape failed because of the applied limit of 8," the next action changes. You get the grounds to decide whether to restart the process, investigate the network, or fix the instrumentation design. Even if a single observation agrees, you have to look at the check time and the most recent scrape time together so that you do not mistake an earlier result for the result of a new change.

What you will do in the next lab

Starting from the normal baseline, you preserve the raw response, the applied configuration, and the collection result. You increase the number of labels to observe the limit rejection, and briefly raise the limit to confirm that the same raw data actually comes in. After that, you compare two wrong fixes with a fix in the source instrumentation. Every service binds only to the personal VM's loopback, and you do not touch anyone else's monitoring server or production configuration. When the lab ends, the VM and TSDB are reclaimed, so keep any observations you need separately before it expires.

Official documentation