PCA — Prometheus Certified Associate
You Only Know a Query by Running It
In one line
PromQL does not raise an error when it is wrong; it gives an empty result. So you can only know whether a query is right by running it on real metrics.
Why it has to be a real Prometheus
In the earlier modules you wrote several PromQL queries. But the place where those labs run has no Prometheus itself, so there was no way to check whether a query gave the answer you wanted.
PromQL often produces a silently empty result even when the syntax is right. It is an empty result rather than an error, so it is hard to notice.
- If a range selector is shorter than the scrape interval,
rate()does not have enough points to compute - If you get one label name wrong, there are no matching time series
- A recording rule computes only from after it is registered, so the past range is empty
What up == 0 cannot catch
This is the most valuable trap. up == 0 is true only when the target exists and does not respond.
If a Pod disappears entirely, the target drops out of service discovery and the up time series itself goes away. Something that does not exist is not 0, so the alert does not trigger.
"Everything died and not a single alert came" comes from here.
up == 0 대상이 있는데 응답이 없다
absent(up{job="x"}) 그 job 의 대상이 아예 없다
count(up{job="x"}) < 3 예상보다 적다 ← 실무에서 가장 쓸모 있다
Where configuration does not take effect
A ConfigMap volume is synchronized periodically by the kubelet (about 60 seconds). If you hit reload before propagation, it rereads the old configuration, and reload returns 200. So you believe the configuration is wrong and fix the wrong place.
And if you mount with subPath, updates are never propagated.
Rules that make alerts useful
When alerts multiply, people stop looking at them, and then it is the same as having none. If you follow four rules, most of them get filtered out.
Attach them to symptoms, not causes. "CPU over 80%" is no reason to wake someone. "Payment error rate over 5%" is. Keep cause metrics on dashboards, and attach alerts only to what users experience.
Always add for. Waking people for a momentary spike loses their trust.
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) > 0.05
for: 10m # 10분 동안 계속 참일 때만
labels: {severity: page}
annotations:
summary: "5xx 비율 {{ $value | humanizePercentage }}"
runbook: "https://…/runbooks/high-error-rate"
Attach a runbook link. An alert that does not say what a person woken at 3 a.m. should look at only creates anxiety.
Attach them to the burn rate. If the SLO is 99.9%, the monthly error budget is 43 minutes. If you alert on "when will we use up the budget at the current speed," a short, severe one and a long, weak one are caught by a single rule.
빠른 소진: 1시간 창에서 예산의 2% 이상 → 즉시 호출
느린 소진: 6시간 창에서 예산의 5% 이상 → 티켓
Metric cardinality kills Prometheus
One label value is one time series. If you put user_id in as a label, as many time series appear as there are users,
and that much memory is needed.
# 지금 무엇이 시계열을 많이 쓰나
topk(10, count by (__name__)({__name__=~".+"}))
# 특정 지표의 라벨별 카디널리티
count(count by (route) (http_request_duration_seconds_count))
The top metrics and labels also appear on Prometheus's /status/tsdb screen. Beyond 1 million
time series, memory grows by several GB, and on a restart the WAL replay takes long, so metrics
look empty for several minutes.
Things whose values can grow without bound (request IDs, emails, full URLs) go into logs or traces, not metrics. This distinction is half of observability design.
What really matters in practice
Alert with count(...) < N, not up == 0. A vanished target is not 0 but nonexistent, so up == 0 stays silent in the biggest incidents. It is very valuable to have one count alert for each job whose expected count you know.
After you fix the configuration, do not trust the 200 from reload; read back the value that took effect. ConfigMap propagation takes up to 60 seconds and a subPath mount is never updated at all. Checking the configuration currently running with /api/v1/status/config is the only reliable way.
A recording rule computes from the moment it is registered. If you move a dashboard over to rule-based queries, the past range looks empty. When migrating, you have to keep the two expressions side by side for a while and compare the overlapping range.
In the next lab, you check all of these yourself on a real Prometheus.