PCA — Prometheus Certified Associate
Reading cAdvisor and node exporter with Real Captured Data
Goal
You read the numbers of cAdvisor and the node exporter with real measured metrics taken from the 7-node cluster that runs this site. When you are done, looking at container CPU, memory, and throttle metrics, you can answer "is this a lot?" with grounds.
Why it matters
Even if you know all the PromQL syntax, you cannot interpret a number if you do not know what the metric counts. In a cumulative counter's magnitude, a long-lived container always wins, working set and usage count different things, and the CPU limit is told by kube-state-metrics, not cAdvisor. If you do not know these three, dashboards are pretty but the verdicts are wrong every time. The data you handle here comes from a real cluster, so the values are not tidy — throttle is 0.19%, and a container with no limit has no metrics at all. That is how the field is.
About the data
- The window is a fixed 3 hours. Check it with
lab-realdata. - You throw queries with
promq-at "<PromQL>". This helper asks at a fixed time inside the window. If you just usepromq, it is the present time, so the result is empty. - The original text of the exporter exposition format is in
/opt/lab/realdata/expo/. - The source, the collection time, and what was touched up are written in
/opt/lab/realdata/README.md. Internal IPs and host names were not erased — this is this site's actual cluster.
Steps
- Check the window with
lab-realdata, and in themonitoringnamespace, write a query that counts how many Pods emitcontainer_cpu_usage_seconds_totalin/root/pca-exporters/01-count.promql(it is the number of Pods, not the number of series). - In
/root/pca-exporters/02-cpu-rate.promql, write a query that gives how many cores theprometheuscontainer inmonitoringis using right now. - Write a query that gives the maximum CFS throttle ratio (
throttled_periods / periods, window[1h]) in/root/pca-exporters/03-throttle.promql, and write the name of the container with the highest ratio on one line in/root/pca-exporters/03-throttle.txt. - For the
grafanacontainer inmonitoring, write a query that gives the difference betweencontainer_memory_usage_bytesandcontainer_memory_working_set_bytesin/root/pca-exporters/04-inactive.promql, and write the name of the metric that gives the same value as that difference in/root/pca-exporters/04-inactive.txt. - Read the QoS class of the
ttscontainer in theainamespace from theidlabel and write it on one line in/root/pca-exporters/05-qos.txt. - Find at least three metric names that are in
/opt/lab/realdata/expo/cadvisor-nuc1.txtbut not in Prometheus, and write them one per line in/root/pca-exporters/06-dropped.txt. - In the
ainamespace, write a query that gives the highest CPU utilization against the limit in/root/pca-exporters/07-limit-join.promql. - Write a query that gives the CPU utilization (0–1) of the node
192.168.219.120:9100in/root/pca-exporters/08-node-cpu.promql. - Write a query that gives the difference between
MemAvailableandMemFreeof the same node in/root/pca-exporters/09-mem-available.promql.
Notes
- Put the query inside the quotes, like
promq-at "count(up)". Throw a query you wrote in a file withpromq-at "$(cat <파일>)"(the placeholder is the file). - Common mistake 1: if the result is empty, suspect the time, not the query.
promqasks at the present time, and the real-measurement data exists only in a fixed window in the past. - Common mistake 2: if you just divide or subtract two metrics with different label sets, the result is empty.
Match them with
on(...)or narrow to the same labels.
Checking the loaded real-measurement data and counting Pods
lab-realdata tells you the window of times you can query. Throw queries with promq-at — if you throw them at the present time, it is outside the window and an empty result comes back. If you just use count(), you get the number of series, that is, the number of containers. One Pod can have several containers, so you have to group by pod first and then count to get the number of Pods.
Measuring the slope of a cumulative counter
container_cpu_usage_seconds_total is a value that keeps adding up the CPU seconds used since birth. How many cores are in use now comes out only if you measure the slope with rate. As long as it is inside the window, any window from [5m] to [30m] is fine. Narrow it to the single prometheus container, not the whole namespace.
Judging exceeding the limit by the throttle ratio
throttled_periods is a count, so that value alone cannot tell you how serious it is. Divide by the periods of the same window to get a ratio. Find the highest container with topk(1, ...) and read the container label of the result. A container with no limit has no such metric at all.
Finding out what the difference between usage and working set is
Narrow the two metrics to the same labels and subtract. If the label sets differ, the subtraction result comes out empty. There is a separate metric that gives exactly the same value as that difference — browse the metrics of the grafana container that start with container_memory_ using promq-at.
Reading the QoS class from the cgroup path
The id label of container_cpu_usage_seconds_total is the cgroup path itself. Look for the kubepods-besteffort and kubepods-burstable segments. A Guaranteed Pod has no such segment at all. Look at the tts container in the ai namespace.
Finding metrics that are in the original but not in Prometheus
/opt/lab/realdata/expo/cadvisor-nuc1.txt is the original text the kubelet actually emitted. Throw the metric names in it one by one with promq-at. The ones that come back empty are the metrics the stack's metric_relabel dropped. Find at least three and write them down.
Joining with kube-state-metrics to get utilization against the limit
The limit is not on the cAdvisor side, so you take it from kube_pod_container_resource_limits. The two metrics have different label sets, so if you just divide, the result is empty. Reduce both sides to the same labels with sum by (pod,container) or match them with on(pod,container). Do not forget to narrow to resource="cpu" — if the memory limit gets mixed in, the value becomes meaningless.
Getting node CPU utilization from the cumulative time per mode
node_cpu_seconds_total is one series for each core × each mode. If you average the rate of the idle mode over the number of cores and subtract it from 1, you get a utilization between 0 and 1. If you use sum, it gets as large as the number of cores. The node is 192.168.219.120:9100.
Measuring the difference between MemFree and MemAvailable
Narrow the two metrics to the same instance and subtract. This difference is 'memory that the cache is using now but that can be handed over if needed.' The node is the same 192.168.219.120:9100 as in the previous step.