TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Green status, but 77 requests vanished from the metric

Continue in TT Lab

Goal

You compare sample limits, label dropping, metric exclusion, and source instrumentation changes on a real Prometheus and explain whether information was preserved.

Why it matters

The business can be normal while collection fails, and with up=1 the data you need can still disappear. It uses only Prometheus 3.14.0 inside a personal VM, a synthetic exporter, and an independent control target. It is a small experiment of at most 13 samples, not a large-scale load, OOM, or performance test. Do not put in real personal information. The services listen only on loopback and run as nobody. It does not change the production LabHub or any external Prometheus. Unlike the existing Kubernetes labs of the PCA, this lab uses the scraper and TSDB of real local processes. It is a 55-minute lab, so extend it before it expires if you need to. When the session ends, the VM, TSDB, and student files are reclaimed.

Prepared environment

Prometheus is at http://127.0.0.1:9098, the exporter at http://127.0.0.1:9911, and the control target at 9912. The configuration is /etc/pca-cardinality/prometheus.json, which is a JSON representation of the YAML that Prometheus reads. The exporter mode is /etc/pca-cardinality/exporter.json, and the student answers and observations are in /root/pca-cardinality. The initial observation is initial.json. Do not restart system services or erase the TSDB. Compare the observed MainPID and InvocationID against the baseline to check that you did not hide the problem with a restart.

python3 /opt/fixtures/pca_cardinality_lab.py observe reads the current configuration, the raw response, and the queries. complete N checks the answer you wrote and preserves the limited change and the actual observation. 2 is increasing the synthetic users, 3 is raising the limit, 4 is dropping the distinguishing label, 5 is excluding the request metric, and 6 is source aggregation and returning the limit. 1, 7, and 8 are observation steps. The key=value in the instructions is an explanation, so write the files to match the JSON format examples. The observation has both the actual loaded configuration and the config declaration. Read the original text and the currently applied result side by side. solve N is the same as viewing the answer, and creates only the answers that do not exist. You must fix existing wrong or partial answers yourself. prepare N prepares only the earlier steps. It does not create the current answer. grade N only reads.

Steps

  1. From the actual baseline in initial.json, confirm a raw response of 5 samples, 4 request time series, and business HTTP 200. In baseline.json, write the string job=pca-cardinality and the integers raw_samples=5, current_series=4, and business_http=200, and run complete 1. The independent pca-sentinel target must have a value of 7 and up of 1.
  2. In overflow.json, write the integers sample_limit=8 and raw_samples=13 and the booleans whole_scrape_failed=true and business_failed=false, and run complete 2. The helper increases the synthetic users from 4 to 12. You observe a new scrape in which the raw HTTP is 200 yet that job has up=0 and a sample limit error. Do not report that only 8 partially succeeded.
  3. In budget.json, write the integers sample_limit=16, current_series=12, and total=78 and the boolean fixes_instrumentation=false, and run complete 3. The helper briefly raises sample_limit and performs a promtool check and reload. Compare the actual applied configuration, the 13 raw samples, the 12 current request time series, and the total 78.
  4. In collision.json, write the string dropped_label=user_id, the boolean aggregates_values=false, and the integers reported_total=1 and up=1, and run complete 4. You look at the three samples when only the distinguishing label is erased in metric relabeling. In this fixed input, the raw total is 78 but the current query is 1. Do not regard labeldrop as summation or guess that up must be 0.
  5. In drop.json, write the string dropped_metric=pca_checkout_requests_total and the booleans requests_missing=true and up_implies_complete=false, and run complete 5. You read the actual configuration that excludes the request metric family. Of the 13 raw responses, only one business gauge remains and up=1, but the request query is an empty vector. Distinguish this from the value 0.
  6. In instrument.json, write the string mode=aggregate, the integers sample_limit=8, current_series=1, and total=78, and the boolean user_id_removed_at_source=true, and run complete 6. The exporter exposes the route total 78 at the source and the temporary filter is removed. You check the three samples in which the raw response is 2 samples from the start and the total is preserved. The run identifiers of the two services must be the same as at the start.
  7. Read the number in result.snapshot.at of observation-3.json as is. In history.json, write historical_time=that number and the integers historical_series=12 and current_series=1 and the boolean deletes_history=false, and run complete 7. Compare the actual query at the past time, count=12 and sum=78, with the current count=1. Do not prove past samples with only the series metadata list.
  8. In report.json, write the integers sample_limit=8 and semantic_total=78 and the booleans unsafe_identifiers=false and zero_loss_from_up=false, and run complete 8. You check current information preservation, the independent control group, no service restart, and the past sample lookup. Do not generalize the lab limit of 8 into a recommended value for every service or claim that personal information was deleted automatically.

Notes and limitations

Fixes are made only within the declaration scope the helper has reviewed. If a pending journal remains, the same request is not repeated automatically. If the declaration change and reload completed and only the observation failed, check that the stored change and the declaration are the same and retry only the observation. If you change completed answers, observations, or journals directly, later steps fail too. The hash is a device for detecting accidental overwrites and is not a security guarantee that prevents every forgery by the same VM root. Grading has a 60-second budget and preparing earlier steps has a 90-second budget. The total of 1 seen in the label drop is an observation of a fixed version and input. It is not a general contract that the first value always remains. An empty vector is different from the value 0. The current aggregation is not a deletion of past samples from the TSDB. The administrative data deletion API is not turned on. The raw counter values are synthetic cumulative counts and not a per-second rate. Also record the detailed questions you lose when you reduce the metric dimensions. Three short observations do not guarantee long-term performance or loss rate. 8 and 16 are lab values, not recommended production limits. Official configuration documentation · HTTP API

Reading the baseline of raw samples and current time series

From the actual baseline in initial.json, confirm a raw response of 5 samples, 4 request time series, and business HTTP 200. In baseline.json, write the string job=pca-cardinality and the integers raw_samples=5, current_series=4, and business_http=200, and run complete 1. The independent pca-sentinel target must have a value of 7 and up of 1.

One metric name can have several label sets.

Creating a state where the business is normal but only collection fails

In overflow.json, write the integers sample_limit=8 and raw_samples=13 and the booleans whole_scrape_failed=true and business_failed=false, and run complete 2. The helper increases the synthetic users from 4 to 12. You observe a new scrape in which the raw HTTP is 200 yet that job has up=0 and a sample limit error. Do not report that only 8 partially succeeded.

Separate the business HTTP from that job's scrape result.

Telling a raised limit from an instrumentation improvement

In budget.json, write the integers sample_limit=16, current_series=12, and total=78 and the boolean fixes_instrumentation=false, and run complete 3. The helper briefly raises sample_limit and performs a promtool check and reload. Compare the actual applied configuration, the 13 raw samples, the 12 current request time series, and the total 78.

See whether the current number of time series also went down when you raised the limit.

Verifying the total when a distinguishing label is erased

In collision.json, write the string dropped_label=user_id, the boolean aggregates_values=false, and the integers reported_total=1 and up=1, and run complete 4. You look at the three samples when only the distinguishing label is erased in metric relabeling. In this fixed input, the raw total is 78 but the current query is 1. Do not regard labeldrop as summation or guess that up must be 0.

Erasing a label and adding values are different operations.

Diagnosing the green light obtained by dropping a metric

In drop.json, write the string dropped_metric=pca_checkout_requests_total and the booleans requests_missing=true and up_implies_complete=false, and run complete 5. You read the actual configuration that excludes the request metric family. Of the 13 raw responses, only one business gauge remains and up=1, but the request query is an empty vector. Distinguish this from the value 0.

Distinguish an empty vector from a sample whose value is 0.

Reducing while preserving meaning in source instrumentation

In instrument.json, write the string mode=aggregate, the integers sample_limit=8, current_series=1, and total=78, and the boolean user_id_removed_at_source=true, and run complete 6. The exporter exposes the route total 78 at the source and the temporary filter is removed. You check the three samples in which the raw response is 2 samples from the start and the total is preserved. The run identifiers of the two services must be the same as at the start.

Compare starting from the raw exporter response, not just the number after the filter.

Telling the current query from past samples

Read the number in result.snapshot.at of observation-3.json as is. In history.json, write historical_time=that number and the integers historical_series=12 and current_series=1 and the boolean deletes_history=false, and run complete 7. Compare the actual query at the past time, count=12 and sum=78, with the current count=1. Do not prove past samples with only the series metadata list.

You need a query that specifies the evaluation time of that moment with the time parameter.

Reporting the scope of the collection budget and information preservation

In report.json, write the integers sample_limit=8 and semantic_total=78 and the booleans unsafe_identifiers=false and zero_loss_from_up=false, and run complete 8. You check current information preservation, the independent control group, no service restart, and the past sample lookup. Do not generalize the lab limit of 8 into a recommended value for every service or claim that personal information was deleted automatically.

up does not guarantee that the meaning of every needed metric is preserved.