PCA — Prometheus Certified Associate
Treating Time as a Value — timestamp(), Derivatives, Instrumentation APIs, Spans
In one line
timestamp() pulls out the time of a sample and time() pulls out the evaluation time as a value, so you can compute "since when" by subtraction. deriv() and predict_linear() read trends from a gauge's slope, a client library registers Counter, Gauge, and Histogram in a registry and exposes them at /metrics, and a trace is a tree of spans that share a trace_id and are linked by parent_id.
Why this was needed
"When was this process last restarted," "how long since it was deployed," and "when will the disk fill" can all be answered only by computing times. The Prometheus instrumentation guidelines set one principle here. It is to export the Unix time at which an event happened, not the elapsed time. If the application updates "N seconds since the last success" itself, the value stops when the update logic stops. If you export the time, time() - my_timestamp_metric always gives the correct elapsed time, and you are free of the problem of the update logic stopping.
The PCA's PromQL domain (28%) has a "timestamp metrics" item, the Instrumentation domain (16%) has "client libraries," and Observability Concepts (18%) has "tracing and spans." All three were covered only briefly in the earlier modules, so here we sort them out in one place.
How it works
time(), timestamp(), and the start time metric
time() returns the seconds since 1970-01-01 UTC, and the documentation stresses that this is not the current time but the time at which the expression is evaluated. It means that when you graph a past range, each point gets that point's time. timestamp(v) returns, in seconds, the time at which each sample of an instant vector was recorded, and treats float and histogram samples the same way.
process_start_time_seconds, which client libraries export as a standard, is "the Unix time (in seconds) when the process started." With this one metric you can catch restarts.
# 지금 기준으로 프로세스가 동작한 시간(초)
time() - process_start_time_seconds
# 최근 1시간 안에 시작 시각이 바뀐(=재시작한) 인스턴스
changes(process_start_time_seconds[1h]) > 0
# 스크레이프가 얼마나 오래됐나 — 샘플 시각과 평가 시각의 차이
time() - timestamp(up)
changes() counts how many times the value changed within the range, so if the start time changes, it reads as a restart. For a counter, resets() counts the decreases and serves the same purpose. The last expression can be read only if you understand staleness. The query time is decided independently of the actual samples, and for each series Prometheus uses as the value at that time the newest sample within the lookback period (5 minutes by default, adjustable with --query.lookback-delta). When a target disappears, the series is soon marked stale and drops out of the result. That is why, if the sample time seen through timestamp() lags the evaluation time by several minutes, it is a sign that scrapes are being delayed.
deriv() and predict_linear() — derivatives for gauges only
If rate() is the per-second growth rate of a counter, deriv(v range-vector) is the per-second derivative obtained by simple linear regression, and predict_linear(v, t) uses the same regression to predict the value t seconds later. The documentation says of both functions that "they should be used only with gauges and work only on float samples," and they compute only if there are at least two float samples in the range. If +Inf or -Inf is mixed in, the result is NaN. When you need the difference of a gauge, use delta(), which extrapolates the difference between the first and last values over the whole range, so a fraction can come out even from integer samples. The instrumentation guidelines state explicitly: do not use rate() on a gauge.
# 4시간 추세로 볼 때 6시간 뒤 남은 디스크가 0 이하가 되는가
predict_linear(node_filesystem_avail_bytes{mountpoint="/data"}[4h], 6 * 3600) < 0
Client libraries — register, expose, and get read at scrape time
The official libraries are Go, Java/Scala, Node.js, Python, Ruby, and Rust, and a library exports the current state of all the metrics it is tracking at scrape time. It is a structure where values are read, not pushed out. Of the four core types, a Counter only goes up and returns to 0 on restart (do not use it for values that can go down), a Gauge goes up and down, and a Histogram counts observations into buckets and exposes them as three series (for the classic histogram): _bucket{le}, _sum, and _count.
from prometheus_client import Counter, Histogram, start_http_server
REQS = Counter("requests_total", "Total requests",
labelnames=["method"], namespace="myapp")
LAT = Histogram("request_duration_seconds", "HTTP request latency",
labelnames=["method", "endpoint"], namespace="myapp",
buckets=[.01, .05, .1, .25, .5, 1, 2.5, 5])
def handle(method, endpoint):
REQS.labels(method=method).inc()
with LAT.labels(method=method, endpoint=endpoint).time():
...
start_http_server(8000) # /metrics 노출
In Python, a Counter strips the trailing _total from its name and adds it back when exposing (because OpenMetrics requires _total). namespace, subsystem, and name are joined with underscores to form the full name, and registry is the default REGISTRY, with None meaning it is not registered (for test code). The le of a Histogram is a reserved label and cannot be used as a label name, buckets must be in ascending order, and +Inf is always added automatically. The default buckets are .005 .01 .025 .05 .075 .1 .25 .5 .75 1 2.5 5 7.5 10.
In Go, you create a registry with prometheus.NewRegistry(), register the runtime and process collectors with reg.MustRegister(collectors.NewGoCollector(), collectors.NewProcessCollector(...)), create metrics with promauto.With(reg).NewCounter(prometheus.CounterOpts{Name: ..., Help: ...}), and hook promhttp.HandlerFor(reg, ...) to /metrics. If you leave Buckets of HistogramOpts empty, DefBuckets (.005 .01 .025 .05 .1 .25 .5 1 2.5 5 10) is used, and the documentation says this value is tuned to measure network service response times broadly, so most people will have to define buckets that suit their own purpose.
Bucket design
In a classic histogram, the buckets are fixed at instrumentation time, one series is created per bucket (even if empty), and changing them later causes great confusion because different layouts cannot be aggregated together. The guidelines say to choose buckets to fit the expected range of values and the queries you want to run. For example, if there is an SLO of "95% of requests within 300ms," you must put a boundary at 0.3 so that _bucket{le="0.3"} gives an exact ratio. A quantile is computed on the server with histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m]))), and the estimation error is bounded by the width of the bucket the quantile falls in. A native histogram (supported in Go and Java) does not choose buckets but only sets a resolution, so the documentation recommends it where possible. Another guideline is do not make metrics that do not exist. If a series does not exist until an event occurs, queries become hard, so export 0 in advance, and for metrics without labels most libraries output 0 automatically.
Traces and spans
A trace is the path a request took through the application, and a span is a unit of work within it. A span holds a name, a parent span ID (empty for the root), start and end times, a span context (trace ID, span ID, trace flags, trace state), attributes, events, links, and a status. Spans of the same trace share the same trace_id, and a child's parent_id equals its parent's span_id. These two fields alone build the tree. The OpenTelemetry documentation likens a span to "a structured log with context, correlation, and hierarchy."
Crossing a service boundary requires context propagation. The default propagator is W3C TraceContext, which carries it in the traceparent header in the format <version>-<trace-id>-<parent-id>-<trace-flags> (for example, 00-a0892f3577b34da6a3ce929d0e0e4736-f03067aa0ba902b7-01), and the receiving side extracts it and makes it the parent of a new span. Propagation is usually done automatically by the instrumentation library. The link between metrics and traces is the exemplar. As in Python's observe(0.43, exemplar={"trace_id": "..."}), you can attach a trace_id to an observation, and it is exposed only in the OpenMetrics format.
What it looks like in the field
There is the case of "the memory leak alert comes every dawn but it is normal in the morning." If you apply predict_linear with a 30-minute range, it extrapolates several hours ahead from the slope during just the 30 minutes the nightly batch runs. Setting the range longer than the batch cycle or adding a for duration is the response that fits the guidelines.
In restart detection, people sometimes use up == 0 instead of changes(process_start_time_seconds[1h]) and miss it. If the restart finishes within the scrape interval, up never becomes 0. The start time metric is a value the process exports itself, so even a short restart leaves a changed value.
What to check in the next quiz
It asks about the exact meaning of the time time() returns, expressions built from timestamp() and the start time metric, the lookback default, the conditions deriv/predict_linear require, the registration method and default buckets of the Python and Go instrumentation APIs, the bucket design guidelines, and the required elements of a span and the traceparent header format. References: PromQL Functions, Querying basics — Staleness, Instrumentation practices, Histograms and summaries, client_python Histogram, Instrumenting a Go application, OpenTelemetry Traces, Context propagation.