The Three Signals Are an Investigation Order, Not Substitutes
In one line
Traces answer "where", metrics answer "when and how much", and logs answer "why". So the investigation order is fixed — metrics → traces → logs.
Why this matters
When an outage hits, people usually open the logs first. And thirty minutes later they are still reading logs. Logs are the most detailed but the most expensive, and if you open them without knowing what to look for, they are just a sea of text.
There are three reasons metrics come first. They are cheap, always on, and low in cardinality. Because they satisfy all three at once, metrics are the only signal that can serve as the basis for alerts. Traces are sampled, so they cannot guarantee "it is getting worse right now", and logs are too voluminous for real-time aggregation costs to be affordable.
The three signals do not replace one another. There are just three bridges connecting them. From metrics to traces you cross over with an exemplar, from traces to logs with the trace_id field, and from logs back to metrics with the event field. Without these bridges, even if you buy all three tools, the investigation still depends on human intuition.
How it works
The real decision in metrics is the choice of type. The choice of type is not a matter of taste. If you choose wrongly, the calculation you want to do later becomes impossible in principle.
- Counter: increases monotonically and goes back to 0 when restarted. The value itself is meaningless; it becomes meaningful only when viewed as a rate.
- Gauge: goes up and down. It represents the current state.
- Histogram: a set of bucket counters. Use it for values whose quantiles you want to know.
- Summary: computes percentiles inside the instance and exports them. So there is no way to combine multiple instances. It cannot be used as a service level indicator.
Why a gauge will not do for latency is confirmed by arithmetic. If the scrape interval is 15 seconds and there are 500 requests per second, a gauge stores the value of only 1 request out of 7,500. The rest become as if they never existed. You cannot compute p99 from this time series, and reprocessing it later does not restore it.
What it looks like in the field
You often come across dashboards where the request rate spikes or drops after every deployment. The cause is almost always the same: rate was applied outside sum. rate treats every moment when a value becomes smaller than the previous sample as a counter reset and adds the value just before the drop to the increase, but if you take the sum first, this correction is applied to the total rather than to each Pod's counter. If one Pod's restart makes the total drop, the whole total is treated as reset and a false spike appears; and if the total does not drop because it is masked by the growth of other Pods, the reset is missed and the result comes out lower by the accumulated value of the restarted Pod. The correct form is sum(rate(x[5m])) by (route), and rate(sum(x)[5m:]) is wrong.
Another is a silent wrong answer. If you leave out by (le) in histogram_quantile, it returns a wrong number without an error. Nobody sees an exception, so that panel stays on the dashboard for months.
Cardinality sets the budget
The hardest mistake to reverse in metric design is putting a label with infinitely many values. The number of time series grows as the product of the label value combinations. If there are three labels with 10, 5, and 4 values each, that is 200 time series, but the moment you add a user ID, it is multiplied by the number of users. With 100,000 users, that is 20 million.
The especially dangerous labels are well known. User IDs, request IDs, session IDs, emails, IP addresses, and paths put in without processing. If you use a path with an identifier embedded, like /orders/8213, as a label, one time series is created per order. The correct form is to normalize it into a route pattern like /orders/:id, and this normalization is a value the framework already knows when it routes, so you can usually get it for free.
The information you cut off this way is not thrown away but moved to another signal. That is why the three signals are kept separate.
| What you want to know | Where to put it | Why |
|---|---|---|
| How slow this route is | Metrics | The kinds of values are few and it must always be on |
| Where this one slow request spent its time | Traces | A request-level identifier attaches naturally |
| Who the user who made that request is | Logs | There is no limit on count and you find it by search |
Histograms need special care. One bucket is one time series, so if you attach 200 label combinations to a histogram with 20 buckets, that one metric creates over 4,000 time series. So use histograms only for metrics that really need quantiles, and set bucket boundaries tightly near the service level objective and sparsely elsewhere. If the objective is 300ms and the buckets jump from 100ms straight to 1 second, you cannot use that histogram to decide whether p99 exceeded 300ms.
Symptoms usually appear first not on the dashboard but on the storage side. Prometheus memory keeps growing, scrapes do not finish in time, and queries slow down. At that point, if you count from the metrics with the most time series using topk(10, count by (__name__)({__name__=~".+"})), the culprit is almost always one or two.
What you will do in the next lab
In the lab environment, a real Prometheus server is running with 12 hours of data in it. It contains two periods in which errors spike, one period in which only the latency tail spikes, and a disk that steadily shrinks.
You do not just write queries to a file; you throw them directly with promq and read the numbers that come back. Whether the queries you wrote find those three events tells you whether you are right — PromQL is hard not because of the syntax, but because even a wrong query returns plausible-looking numbers without an error.