TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Scrape Configuration and Relabelling

Continue in TT Lab

Goal

You write a Prometheus scrape configuration yourself from start to finish, and check how that configuration picks the Pods of a real cluster. When you are done, you can look at someone else's prometheus.yml and say which targets survive and which metrics are dropped.

Why it matters

The scrape configuration is the only place in Prometheus that decides "what to look at." All the queries and alerts after it deal only with data that survived here. Relabeling feels hard not because of the syntax but because the two stages operate at different times. relabel_configs handles the target list before the scrape, and metric_relabel_configs handles the metrics that came in after the scrape. A target dropped in the earlier stage costs 0 because no request goes out at all, while a metric dropped in the later stage has already paid the network and parsing cost. This difference decides whether Prometheus lives or dies in a large cluster. Guardrails are the same story. Without sample_limit, a single service whose cardinality has blown up halts all monitoring.

Steps

  1. Create /root/pca-scrape/prometheus.yml and write the global block. scrape_interval: 15s, evaluation_interval: 30s, scrape_timeout: 10s, and under external_labels, cluster: homelab.
  2. In the first item of the scrape_configs list, create job_name: prometheus-self with metrics_path: /metrics, and put a single localhost:9090 in the targets of static_configs.
  3. In the second item, create job_name: checkout-api and put in scrape_interval: 30s, scrape_timeout: 20s, sample_limit: 20000, label_limit: 24, and label_value_length_limit: 256.
  4. In the third item, create job_name: kubernetes-pods, specify role: pod in kubernetes_sd_configs, and then put only pca-scrape in namespaces.names to narrow the watch scope.
  5. In the third job's relabel_configs, put three rules. (a) action: keep, source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape], regex: "true" (b) action: replace, source_labels: [__meta_kubernetes_pod_ip, __meta_kubernetes_pod_annotation_prometheus_io_port], target_label: __address__, replacement: '$1:$2' (c) action: replace, source_labels: [__meta_kubernetes_namespace], target_label: namespace.
  6. To the same job, add the rule action: labelmap, regex: __meta_kubernetes_pod_label_(.+), and put two rules in metric_relabel_configs. (a) action: drop, source_labels: [__name__], regex: 'go_gc_duration_seconds.*|python_gc_.*' (b) action: replace, source_labels: [__name__], regex: 'http_requests_total', target_label: user_id, replacement: ''.
  7. Create the namespace pca-scrape, and in it create the ConfigMap prometheus-config. The key name is prometheus.yml and the value is the contents of the file you just wrote.
  8. In the namespace pca-scrape, create a Pod checkout-api. Label app: checkout-api, annotations prometheus.io/scrape: "true", prometheus.io/port: "8080", prometheus.io/path: "/metrics", and the container port is containerPort: 8080 with the name metrics.

Notes

Writing the global block

Create /root/pca-scrape/prometheus.yml and write the global block. scrape_interval: 15s, evaluation_interval: 30s, scrape_timeout: 10s, and under external_labels, cluster: homelab.

Under global at the top level of prometheus.yml, put scrape_interval, evaluation_interval, and scrape_timeout. external_labels are labels attached to every time series this Prometheus emits, so they are used to tell instances apart in federation or remote storage.

Adding a static_configs job

In the first item of the scrape_configs list, create job_name: prometheus-self with metrics_path: /metrics, and put a single localhost:9090 in the targets of static_configs.

scrape_configs is a list. Put job_name and static_configs in the first item. static_configs is a list of items that each have a targets list, so brackets appear twice. metrics_path has a default, but state it explicitly here.

A job with guardrails

In the second item, create job_name: checkout-api and put in scrape_interval: 30s, scrape_timeout: 20s, sample_limit: 20000, label_limit: 24, and label_value_length_limit: 256.

sample_limit is the upper bound on the number of samples one target gives, and if it is exceeded, that scrape fails entirely. label_limit and label_value_length_limit are the same kind of defense. Remember the constraint that scrape_timeout cannot be larger than scrape_interval.

Finding Pods with kubernetes_sd

In the third item, create job_name: kubernetes-pods, specify role: pod in kubernetes_sd_configs, and then put only pca-scrape in namespaces.names to narrow the watch scope.

kubernetes_sd_configs is also a list. The role is one of node/service/pod/endpoints/endpointslice/ingress, and to scrape a Pod's container port directly, it is pod. If you narrow the watch scope with namespaces.names, the SD load drops greatly on a large cluster.

Choosing with keep and reassembling the address

In the third job's relabel_configs, put three rules. (a) action: keep, source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape], regex: "true" (b) action: replace, source_labels: [__meta_kubernetes_pod_ip, __meta_kubernetes_pod_annotation_prometheus_io_port], target_label: __address__, replacement: '$1:$2' (c) action: replace, source_labels: [__meta_kubernetes_namespace], target_label: namespace.

keep drops the targets that do NOT match the regex. Remember the rule that in annotation meta label names, dots and slashes become underscores. If you use two source labels, the values are joined with a semicolon and enter the regex, and in replacement you reassemble them with capture groups.

labelmap and metric_relabel_configs

To the same job, add the rule action: labelmap, regex: __meta_kubernetes_pod_label_(.+), and put two rules in metric_relabel_configs. (a) action: drop, source_labels: [__name__], regex: 'go_gc_duration_seconds.*|python_gc_.*' (b) action: replace, source_labels: [__name__], regex: 'http_requests_total', target_label: user_id, replacement: ''.

relabel_configs handles targets before the scrape, and metric_relabel_configs handles metrics after the scrape. labeldrop matches only on the label 'name' and erases it from every metric of that job, so to erase it only from a specific metric, put in an empty value with replace. A label with an empty value is the same as a label that does not exist.

Putting the configuration up as a ConfigMap

Create the namespace pca-scrape, and in it create the ConfigMap prometheus-config. The key name is prometheus.yml and the value is the contents of the file you just wrote.

The --from-file of kubectl create configmap can specify the key name in the 키=경로 form (the placeholders are the key and the path). After editing the file, you have to recreate the ConfigMap for it to take effect. In real operation, a config-reloader sidecar detects this update and triggers a reload in Prometheus.

Creating a Pod that the keep rule will let through

In the namespace pca-scrape, create a Pod checkout-api. Label app: checkout-api, annotations prometheus.io/scrape: "true", prometheus.io/port: "8080", prometheus.io/path: "/metrics", and the container port is containerPort: 8080 with the name metrics.

Choose the annotation values while thinking about whether the keep rule and the address reassembly rule you wrote earlier let this Pod through. The value must be exactly equal to the string true, and the port annotation value must match the container port. Also attach one Pod label for the labelmap rule to carry over.