TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Suffixes and Colons Are a Contract, Not a Convention

Continue in TT Lab

In one line

Prometheus's metric naming rules are not a matter of taste but a contract that tools depend on. _total is a declaration that it is a counter, _bucket, _sum, and _count are the three components of a histogram, and colons are allowed only in recording rules. Only when you can tell what something is from its name alone can dashboards and alerts trust each other.

Why this was needed

If you choose the wrong type, the calculation you want to do later becomes fundamentally impossible. The most common incident is recording latency as a gauge.

last_latency = Gauge("http_request_duration_seconds", "요청 처리 시간")
last_latency.set(elapsed)   # 앞선 값들은 전부 사라진다

With a scrape interval of 15 seconds and 500 requests per second, only 1 value out of 7,500 requests is stored. The rest become as if they never existed. You cannot compute p99 from this time series, and reprocessing the data later does not restore it.

How it works

The exposition format is human-readable text. Each line holds the metric name, the labels inside braces, and the value, with # HELP and # TYPE comments attached. The standard path is /metrics, and if you use a different path, you must specify metrics_path in the scrape configuration. OpenMetrics standardizes this format and adds exemplars and created timestamps.

Four types

Type Exposed time series The question it answers
Counter x_total How fast does it grow per second (needs rate)
Gauge x What is the value right now
Histogram x_bucket{le}, x_sum, x_count What is the distribution (quantile estimated on the server)
Summary x{quantile}, x_sum, x_count The quantile within this process

The reason Summary is useless at the service level is one property covered in the earlier module. Quantiles cannot be combined. A Summary emits quantiles the client has already computed, so there is no way to merge the p99 of 40 Pods into one. A Histogram gives up accuracy down to the bucket resolution in exchange for aggregatability. In a distributed system, the latter is almost always the right trade.

Which component a metric comes from is also an exam regular.

A recording rule moves computation from query time to evaluation time. There are four criteria for creating one. When the same expression repeats in three or more places, when a query takes more than 2 seconds, when an alert runs a heavy expression at every evaluation, and when you want to aggregate a high-cardinality source and keep it for a long time.

Names follow the 수준:메트릭:연산 convention (the placeholders are the level, the metric, and the operations). Looking at route:http_requests:rate5m, you can read from the name alone that it is the 5-minute rate of the request count aggregated to the route level. Colons are never used in the names of raw metrics, so from the name alone you can tell "this is a derived time series."

The hierarchy is the key. Once layer 1 has scanned the raw data once, layer 2 references only that result, so the raw scan ends after one time. But within a group, rules are evaluated from top to bottom, so layer 1 must be above layer 2.

What it looks like in the field

I once put in 2,000 recording rules all at once, which created about 1 million time series and made the WAL surge. The calculation before deployment is simple. With 120 routes, creating a rule for each of 5 window types gives 600, and with 50 rules that is 30,000. Review a rule file like code, but its cost is the same as a new metric.

Validation runs in CI. You check syntax with promtool check rules and expected values with promtool test rules. In particular, if you calculate the expected values yourself and write them in the unit tests, CI catches it when someone later "optimizes" a rule and changes its meaning. There are also things promtool cannot catch — circular references between rules must be spotted by a person in review.

What you will do in the next lab

Under /root/pca-rules/, you write a recording rule file and an alerting rule file and build a two-level hierarchy with level:metric:operations naming. Next, in a promtool unit test file, you put the input time series and expected values that you calculate yourself, and finally you move the same content into a PrometheusRule CR that the Prometheus Operator reads.