TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Metrics, Logs and Traces Answer Different Questions

Continue in TT Lab

In one line

Metrics answer when and how much, logs answer why, and traces answer where. The three are not supersets of one another; they are tools that can answer differently shaped questions. The moment you ignore this boundary and attach a request ID to a metric, it is no longer a metric but an event store with a poor compression ratio.

Why this was needed

During incident response, you open a dashboard. It has 40 panels. CPU, memory, thread count, GC count, connection pool size, heap usage, requests per second. Every one has a graph drawn. But there is only one thing you want to know right now: "Are users experiencing failures, and if so, what percentage?" That answer is not in any of the 40 panels.

The problem with metric design is not a lack of data but a mismatch between the questions the collected data can answer and the questions that need answering. So the starting point is not a list of metrics but a list of questions. First write down the five questions the on-call engineer asks at 3 a.m., and build only the metrics that answer those five.

How it works

Question Metrics Logs Traces
Since when has it gotten worse Answers Hard Cannot
What percentage of requests were affected Answers Expensive Cannot
Why did this request fail Cannot Answers Cannot
Which service did a slow request spend its time in Cannot Cannot Answers

The cells metrics cannot fill matter. The moment you add labels to fill those cells with metrics, the cost explodes. The limit of metrics is set by cardinality.

And the concept that comes up most often on the exam is the percentile. Suppose that out of 10,000 requests, 9,900 took 50ms and 100 took 3,000ms. The average is 79.5ms. Not a single request was actually answered in 79.5ms. Worse is the sensitivity. Even if the slow 100 get twice as bad, from 3,000ms to 6,000ms, the average only moves from 79.5 to 109.5, and a 30ms rise does not cross any alert threshold.

Two things follow from this.

First, p99 is not "1% of users." It means 1 request out of 100, so if one screen calls 20 APIs, the probability that the screen hits p99 is 18.2%, and a user who makes 200 requests a day has an 86.6% chance of running into the worst range at least once a day.

Second, percentiles cannot be combined. If server A handled all 9,900 requests in 50ms and server B handled all 100 in 3,000ms, the average of the p99 values is 1,525ms, the request-count-weighted average is 79.5ms, and the true p99 of the two servers lined up as one is 3,000ms. All three numbers differ, and the first two mean nothing. Counters can be added and percentiles cannot — this is exactly why histograms exist.

An SLI is a measured indicator, an SLO is its target value, and an SLA is a legal contract. For a 30-day 99.9% target, the error budget is 0.1%, which is 0.1% of 43,200 minutes, or 43.2 minutes. The burn rate is the multiple of the speed at which that budget is being burned, so if you keep burning at 14.4, the 30-day budget is gone in 50 hours.

What it looks like in the field

The author's seven-node homelab (3 control plane + 4 GPU workers) runs kube-prometheus-stack, and Grafana holds 10.0.0.203 in the MetalLB pool 10.0.0.200–215. If you calculate how expensive one histogram is here, you get a feel for it.

http_request_duration_seconds_bucket
  route 120 × method 5 × le 11 × pod 40 = 264,000 시계열
  여기에 _sum 과 _count: 120 × 5 × 40 × 2  = 48,000
  ------------------------------------------------
  합계 약 312,000 시계열 — 메트릭 하나에서

One metric uses 300,000 time series. At roughly 8KB per series, 10 million series is 80GB. This one multiplication settles why "put everything in first and trim later" is dangerous.

To add one lesson learned on the same cluster — when I installed KubeVirt, every component status was AllComponentsReady, yet the VM would not start. "The status is Ready" and "it actually works" are different claims. Observability must be built not on a Ready label but on the symptoms users experience.

How Prometheus sees the world

When you use Prometheus, there are a few things that make you wonder "why does it behave like this," and most of them come from the pull model and the time series storage structure. Once you know those two, the rest is explained.

When a target disappears, its metrics disappear too. When a Pod dies, that time series no longer comes in. So up == 0 cannot catch a target that has vanished — because up itself does not exist. To catch something disappearing, you use absent() or compare against the expected count from service discovery.

A value is a snapshot at scrape time. If you scrape every 15 seconds, a momentary spike in between is not visible. A counter is cumulative so it does not miss anything, but a gauge keeps only the value at the moment of the scrape. If you need to know the instantaneous maximum, the application has to record the maximum itself in a counter or histogram.

Use rate() only on counters. On a gauge it is meaningless, and conversely, if you just plot a counter you see only a line that keeps climbing. rate automatically corrects for a counter going back to 0 on a restart.

The range vector window should be at least 4 times the scrape interval. With rate(x[1m]) and a 30-second interval, there are only two points, so the value is unstable, and if even one is missed the result comes out empty. Something like [5m] is a safe default.

A single differing label makes a different time series. That is why a deployment that changes labels creates a break in the graph. The old time series ends at that moment and a new one begins. If you group totals on a dashboard with sum by (), this break is less visible.

Storage is local, and keeping data long requires another mechanism. Prometheus itself does not aim at long-term retention or high availability. That role is left to remote write and a separate long-term store.

What to look for in the next check

This module is a concept module, so there is no lab. After checking the boundaries of the three signals and the properties of percentiles with the quiz, in the next module you write prometheus.yml yourself and decide by hand "which targets survive and which metrics are dropped."