PCA — Prometheus Certified Associate
up Is Itself a Signal — the Design of Pull and the TSDB
In one line
Prometheus chose pull not for performance but for observability. If the server goes out to scrape, the fact that "there was no response" becomes data in itself, and that is the up metric. The TSDB writes the samples that come in that way sequentially to the WAL first, gathers them in a 2-hour Head Block, and then compacts them into persistent blocks.
Why this was needed
In a push model, you cannot tell the difference between data not arriving and a process having died. Both look like "nothing." The pull model leaves a record of the server's attempt, so failure becomes a value.
up{job="checkout-api"} == 0
or
absent(up{job="checkout-api"})
It often comes up on the exam that the two conditions are different situations. up == 0 is a state where the target was discovered but the scrape failed, and absent(up{...}) is a state where the target has vanished from discovery altogether and the up time series itself does not exist. If you leave out the latter, the alert quietly dies when a service is deleted.
Of course pull is not a cure-all. A short-lived batch job that ends before Prometheus comes to scrape cannot be caught with pull, so there is an exception channel called the Pushgateway. The point of the design is to keep an exception an exception.
How it works
The storage path has three layers.
| Layer | Location | Nature |
|---|---|---|
| WAL | wal/ segments (128MB by default) |
Sequential writes, for crash recovery |
| Head Block | Memory + chunks_head/ mmap |
The most recent ~2 hours, writable |
| Persistent block | ULID directory | meta.json, index, chunks/, tombstones |
A new sample is written to the WAL first and then attached to a memSeries in the Head Block. When the active chunk reaches 120 samples or 2 hours, it is sealed and moved down to chunks_head/ as an mmap file, and the OS page cache takes over memory management. Compaction is level-based, so several 2-hour blocks are combined into 6 hours, and again into 18 hours. Deletion does not happen immediately; it is marked with a tombstone and actually removed at compaction time.
Compression comes from the Gorilla paper. Timestamps are stored as delta-of-delta, so if the scrapes are regular most of them take 1 bit, and values are XORed with the previous value and the leading and trailing 0 bits are trimmed. The result is an average of 1.37 bytes per sample. If you memorize this number, capacity calculations become easy.
Search is an inverted index. For each label name-value pair, a list of time series IDs (a posting list) is stored sorted, so the intersection of job="prometheus" and instance="localhost:9090" is computed in linear time. This is why label selectors are fast, and why this index is the first thing to get heavy as cardinality grows.
The exam targets three points in configuration.
- A change to the configuration file is not detected automatically. You have to send SIGHUP, or turn on
--web.enable-lifecycleand then POST to/-/reload. scrape_timeoutmust not be larger thanscrape_interval. The default is 10 seconds.- Retention is not in the configuration file but a command-line flag (
--storage.tsdb.retention.time, default 15 days). If you combine it with a size-based one, whichever is reached first wins.
And staleness. When a target disappears, a stale marker (a special NaN) is attached to that time series, and a query looks for the most recent sample within the lookback delta (default 5 minutes), and when it meets a stale marker, it drops that time series from the result. The accurate description is not "it disappears after 5 minutes" but "it drops out immediately when it meets the marker."
What it looks like in the field
I once deployed 2,000 recording rules at once. One rule creates as many time series as there are routes, so the result was about 1 million new time series, and as they all went into the Head Block, the WAL surged. A rule file looks like code, but in cost terms it is the same as adding a new metric. You need the habit of multiplying "120 routes × 5 windows × 50 rules" before deploying.
There is also a way to stop it at collection time. Setting sample_limit keeps a single exploded target from bringing down all of Prometheus. When it triggers, that target's scrape fails entirely, so set the value generously but always set it. Without an upper bound, one service's mistake halts all monitoring.
Also get a sense of capacity. At roughly 8KB per series, 10 million series is 80GB. On the homelab's 4C/14GB mini PC control plane, this number is the limit line.
What you will do in the next lab
You write /root/pca-scrape/prometheus.yml from scratch. You include the global block, a static_configs job, a job that finds Pods with kubernetes_sd, rules that pick targets with relabel and reassemble the address, rules that drop unneeded metrics with metric_relabel, and finally a sample_limit guardrail. At the end, you put that file into the cluster as a ConfigMap and create one Pod that the keep rule you just wrote will actually let through, to check it.