TT Lab
Get started
Learn Learning paths Courses

Where Distributed Tracing Breaks

Make metrics and traces answer the same question

Continue in TT Lab

Goal

You count the request count, error count, and latency distribution from the span dump alone, compare with the metrics-side exposition-format file to find where names and values disagree, and align the instrumentation. You swap samplers to produce, from the same data, different error rates yourself and correct them with the reciprocal of the sampling probability, then build a bridge that picks a representative trace per interval, harden the rules into a file, and apply them to a second service.

Why it matters

Metrics and traces are usually built separately. On the metrics side the framework puts in the route template, and on the tracing side the address is put in by hand, so a metric with nine series faces more than thirty bundles of spans. In that state, if you say "let's open just one slow request from this peak," nobody can answer. Worse is counting ratios from spans — under a common sampling policy that keeps every error and only one in five successes, the error rate counted by spans inflates to nearly four times the real one, and if the sampling probability was not written on the span, there is no way to undo it at all. If you attach the two signals with the same name and the same value and write the probability alongside, counts and ratios are read from metrics and traces go back to their proper place, finding examples. Cumulative versus delta in the metrics SDK is another module's job, and what you build here are an attribute rules file and a checker that compares the two signals.

Steps

  1. Create /root/tp-metrics/collect.py (default dump path /root/tp-metrics/01-spans.jsonl). While processing SHOP from the material tracelab.tp_metrics.traffic in order, create a root span GET <주소> (the placeholder is the address) for each request, and pass five attributes when starting the span — request.id, request.index, http.target (the request's path), http.response.status_code (the request's status), and user.id (the request's user). Call traffic.work(tracer, req) inside the span. Then write four lines with two tab-separated columns in /root/tp-metrics/01-from-spans.tsv — requests<탭><루트 스팬 수>, errors<탭><상태 코드 500 이상인 루트 수>, p50_ms<탭><값>, and p95_ms<탭><값> (the placeholders are a tab, the number of root spans, the number of roots with a status code of 500 or more, and the value). For the quantiles, put the root spans' duration_ms in ascending order, count from 1, and choose the value at position 올림(비율 × 개수) (the ceiling of the ratio times the count), writing it to three decimal places.
  2. The file produced by the metrics pipeline is at /opt/app/tracelab/tp_metrics/metrics/shop-api.prom (Prometheus exposition format). Write three lines with three tab-separated columns in /root/tp-metrics/02-mapping.tsv — the first column is the metric label, in order handler, code, and svc, the second column is the place in the step 1 dump that you mean to map to that label (http.target, http.response.status_code, and service.name), and the third column is yes or no for whether the values of the two match as they are. Then write two lines in /root/tp-metrics/02-gap.txt — after metric_series=, the number of http_requests_total series in that file, and after span_groups=, the number of groups you get when you group the step 1 dump's root spans by the http.target value.
  3. Create /root/tp-metrics/aligned.py (default dump path /root/tp-metrics/03-aligned.jsonl). Process the same traffic as in step 1, but add http.route (the request's route, that is, the route template) to the root span when passing attributes, and name the span GET <경로 틀> as well (the placeholder is the route template). Leave http.target, which is used for investigation, as it is. Use samplers.keep_all() as the sampler. After running it, write one line per series, four tab-separated columns, in /root/tp-metrics/03-joined.tsv — <handler><탭><code><탭><지표 값><탭><그 짝의 루트 스팬 수> (the placeholders are the handler, a tab, the code, a tab, the metric value, a tab, and the number of root spans in that pair), in ascending order of handler and, if the same, ascending order of code. The two numbers must be equal for every series.
  4. Create /root/tp-metrics/sampled.py. Take nth5 or errbias as a command-line argument and use samplers.every_nth(5) and samplers.errors_and_nth(5) as the sampler respectively, and instrument everything else exactly as in step 3. The default dump path is /root/tp-metrics/04-<인자>.jsonl (the placeholder is the argument). Run it twice to make /root/tp-metrics/04-nth5.jsonl and /root/tp-metrics/04-errbias.jsonl, and then, with the step 3 dump as the third, write three lines with four tab-separated columns in /root/tp-metrics/04-rates.tsv — the first column is, in order, full, nth5, and errbias (full is the step 3 dump), followed by <오류 루트 수><탭><전체 루트 수><탭><비율> (the placeholders are the number of error roots, a tab, the total number of roots, a tab, and the ratio), and the ratio is to four decimal places.
  5. The root spans of the two sampled dumps have sampling.probability written on them. The number of requests one span represents is the reciprocal of that value. Write two lines with four tab-separated columns in /root/tp-metrics/05-adjusted.tsv — the first column is, in order, nth5 and errbias, followed by <보정한 오류 수><탭><보정한 전체 수><탭><보정한 비율> (the placeholders are the adjusted error count, a tab, the adjusted total, a tab, and the adjusted ratio), where the first two values are to four decimal places and the ratio is to four decimal places too. Then write two lines in /root/tp-metrics/05-limits.txt — after limit1= and limit2=, what cannot be undone even by the correction, each in at least 40 characters. Compare with the full ratio from step 4 to confirm how far the correction gets you.
  6. Create /root/tp-metrics/bridge.py (default dump path /root/tp-metrics/06-bridge.jsonl). After instrumenting as in step 3 and leaving the dump, read that dump again, pick the single slowest root span for each series, and write them in a table. The table's path is the dump path with .jsonl changed to .tsv (by default /root/tp-metrics/06-bridge.tsv). Each line is four tab-separated columns, <http.route><탭><상태 코드><탭><duration_ms(소수 셋째 자리)><탭><trace_id> (the placeholders are the route, a tab, the status code, a tab, the duration_ms to three decimal places, a tab, and the trace_id), in ascending order of http.route and, if the same, ascending order of status code. Then write two lines, limit1= and limit2=, in /root/tp-metrics/06-limits.txt, each in at least 40 characters, about what this bridge cannot answer.
  7. Write seven lines with three tab-separated columns in /root/tp-metrics/07-contract.tsv. The first column is the span attribute name, in order http.route, http.response.status_code, service.name, http.target, request.id, user.id, and sampling.probability. The second column is the label name that attribute uses in metrics, and - for those not kept in metrics. The third column is both (keep the same value in both signals) or trace-only (keep it only in traces). This table must actually hold when checked against the step 3 dump and /opt/app/tracelab/tp_metrics/metrics/shop-api.prom — attributes marked both must all be on the root spans of the step 3 dump, and attribute names marked trace-only must not appear as labels in the metrics file.
  8. Create /root/tp-metrics/pay.py (default dump path /root/tp-metrics/08-pay.jsonl). Instrument traffic.PAY from the material under the service name pay-api, leaving the attributes of the step 7 rules as they are; the root span name is POST <경로 틀> (the placeholder is the route template), and the sampler is samplers.keep_all(). And create /root/tp-metrics/agree.py — when run as python3 agree.py <스팬덤프> <노출형식파일> (the placeholders are the span dump and the exposition-format file), it compares the metric value and the number of root spans for each series, prints one line mismatch<탭><handler><탭><code><탭><지표 값><탭><스팬 수> (the placeholders are a tab, the handler, a tab, the code, a tab, the metric value, a tab, and the span count) for each series that differs and ends with exit code 1, and if everything is the same it prints one line ok<탭><계열 수> (the placeholders are a tab and the number of series) and ends with 0. Save the output of running the checker on /opt/app/tracelab/tp_metrics/metrics/pay-api.prom in /root/tp-metrics/08-agree.txt.

Notes

Count the request count and latency distribution from the span dump alone

Create /root/tp-metrics/collect.py (default dump path /root/tp-metrics/01-spans.jsonl). While processing SHOP from the material tracelab.tp_metrics.traffic in order, create a root span GET <주소> (the placeholder is the address) for each request, and pass five attributes when starting the span — request.id, request.index, http.target (the request's path), http.response.status_code (the request's status), and user.id (the request's user). Call traffic.work(tracer, req) inside the span. Then write four lines with two tab-separated columns in /root/tp-metrics/01-from-spans.tsv — requests<탭><루트 스팬 수>, errors<탭><상태 코드 500 이상인 루트 수>, p50_ms<탭><값>, and p95_ms<탭><값> (the placeholders are a tab, the number of root spans, the number of roots with a status code of 500 or more, and the value). For the quantiles, put the root spans' duration_ms in ascending order, count from 1, and choose the value at position 올림(비율 × 개수) (the ceiling of the ratio times the count), writing it to three decimal places.

Child spans (db.query) also go into the dump, so when counting requests you must count only the lines with no parent_id. The quantile is the position math.ceil(0.95 * n) - 1 (a Python index counted from 0). Elapsed times vary a little by machine, so the grader rereads your dump and computes the same way to compare — it is not about matching absolute numbers. Delete the file before making the dump again.

Find where the metrics-side names and values disagree with the span attributes

The file produced by the metrics pipeline is at /opt/app/tracelab/tp_metrics/metrics/shop-api.prom (Prometheus exposition format). Write three lines with three tab-separated columns in /root/tp-metrics/02-mapping.tsv — the first column is the metric label, in order handler, code, and svc, the second column is the place in the step 1 dump that you mean to map to that label (http.target, http.response.status_code, and service.name), and the third column is yes or no for whether the values of the two match as they are. Then write two lines in /root/tp-metrics/02-gap.txt — after metric_series=, the number of http_requests_total series in that file, and after span_groups=, the number of groups you get when you group the step 1 dump's root spans by the http.target value.

One line of the exposition format is 이름{라벨="값",...} 값 (the placeholders are the name, the label, the value, and the sample value), and lines starting with # are descriptions. If you put the two numbers side by side, why they cannot be paired is visible at a glance — one side is a number you can count by hand, and the other grows by one for every address. service.name is not in the span's attributes but in resource.

Attach the two signals with the same name and the same value

Create /root/tp-metrics/aligned.py (default dump path /root/tp-metrics/03-aligned.jsonl). Process the same traffic as in step 1, but add http.route (the request's route, that is, the route template) to the root span when passing attributes, and name the span GET <경로 틀> as well (the placeholder is the route template). Leave http.target, which is used for investigation, as it is. Use samplers.keep_all() as the sampler. After running it, write one line per series, four tab-separated columns, in /root/tp-metrics/03-joined.tsv — <handler><탭><code><탭><지표 값><탭><그 짝의 루트 스팬 수> (the placeholders are the handler, a tab, the code, a tab, the metric value, a tab, and the number of root spans in that pair), in ascending order of handler and, if the same, ascending order of code. The two numbers must be equal for every series.

Metrics split series by the labels handler and code, and spans are grouped by the attributes http.route and http.response.status_code. The status code is a string in metrics and an integer in spans, so you must bring them to one side when pairing. If you use keep_all(), sampling.probability is written as 1.0 on every root span, and you will use that value in step 5.

Change the sample and the same data gives a different error rate

Create /root/tp-metrics/sampled.py. Take nth5 or errbias as a command-line argument and use samplers.every_nth(5) and samplers.errors_and_nth(5) as the sampler respectively, and instrument everything else exactly as in step 3. The default dump path is /root/tp-metrics/04-<인자>.jsonl (the placeholder is the argument). Run it twice to make /root/tp-metrics/04-nth5.jsonl and /root/tp-metrics/04-errbias.jsonl, and then, with the step 3 dump as the third, write three lines with four tab-separated columns in /root/tp-metrics/04-rates.tsv — the first column is, in order, full, nth5, and errbias (full is the step 3 dump), followed by <오류 루트 수><탭><전체 루트 수><탭><비율> (the placeholders are the number of error roots, a tab, the total number of roots, a tab, and the ratio), and the ratio is to four decimal places.

The sampler decides by looking at request.index and http.response.status_code, so if you do not pass those two attributes when starting the span, the sampler sees nothing. Only one of the three ratios will stand out sharply — think about which one and why. The dump is appended to, so delete it before running.

Correct with the reciprocal of the sampling probability, recount, and write down what cannot be fixed

The root spans of the two sampled dumps have sampling.probability written on them. The number of requests one span represents is the reciprocal of that value. Write two lines with four tab-separated columns in /root/tp-metrics/05-adjusted.tsv — the first column is, in order, nth5 and errbias, followed by <보정한 오류 수><탭><보정한 전체 수><탭><보정한 비율> (the placeholders are the adjusted error count, a tab, the adjusted total, a tab, and the adjusted ratio), where the first two values are to four decimal places and the ratio is to four decimal places too. Then write two lines in /root/tp-metrics/05-limits.txt — after limit1= and limit2=, what cannot be undone even by the correction, each in at least 40 characters. Compare with the full ratio from step 4 to confirm how far the correction gets you.

The adjusted count is the sum of 1 / sampling.probability over the root spans. Child spans have no probability written, so naturally only the roots are counted. For errbias, the correction brings it very close to the step 4 full ratio, and nth5 was probably not far off to begin with. When thinking about what cannot be undone, recall "combinations that never entered the sample" and "quantiles."

Pick a representative trace per interval and build a bridge to cross over from metrics

Create /root/tp-metrics/bridge.py (default dump path /root/tp-metrics/06-bridge.jsonl). After instrumenting as in step 3 and leaving the dump, read that dump again, pick the single slowest root span for each series, and write them in a table. The table's path is the dump path with .jsonl changed to .tsv (by default /root/tp-metrics/06-bridge.tsv). Each line is four tab-separated columns, <http.route><탭><상태 코드><탭><duration_ms(소수 셋째 자리)><탭><trace_id> (the placeholders are the route, a tab, the status code, a tab, the duration_ms to three decimal places, a tab, and the trace_id), in ascending order of http.route and, if the same, ascending order of status code. Then write two lines, limit1= and limit2=, in /root/tp-metrics/06-limits.txt, each in at least 40 characters, about what this bridge cannot answer.

There is a reason to derive the table's path from the dump path — the grader runs your program once more in its own temporary directory, and if you write the table to a fixed path, you would overwrite the file you submitted. The standard calls the act of choosing a representative an exemplar, but this environment has no Prometheus, so you cannot actually store or query it. When writing the limits, recall "an ordinary request" and "a request left out of the sample."

Harden into a rules file which attributes to keep common to both signals

Write seven lines with three tab-separated columns in /root/tp-metrics/07-contract.tsv. The first column is the span attribute name, in order http.route, http.response.status_code, service.name, http.target, request.id, user.id, and sampling.probability. The second column is the label name that attribute uses in metrics, and - for those not kept in metrics. The third column is both (keep the same value in both signals) or trace-only (keep it only in traces). This table must actually hold when checked against the step 3 dump and /opt/app/tracelab/tp_metrics/metrics/shop-api.prom — attributes marked both must all be on the root spans of the step 3 dump, and attribute names marked trace-only must not appear as labels in the metrics file.

The criterion for separating them is cardinality. In a metric's labels, the number of distinct values multiplies the number of series directly, so if you put in an address or a user identifier, the series explode. Conversely, traces exist to find a single request, so such values must be there. service.name is in the span's resource, yet it is still both — because it is the key that groups the two signals by service.

Apply the rules to a second service and compare with a checker

Create /root/tp-metrics/pay.py (default dump path /root/tp-metrics/08-pay.jsonl). Instrument traffic.PAY from the material under the service name pay-api, leaving the attributes of the step 7 rules as they are; the root span name is POST <경로 틀> (the placeholder is the route template), and the sampler is samplers.keep_all(). And create /root/tp-metrics/agree.py — when run as python3 agree.py <스팬덤프> <노출형식파일> (the placeholders are the span dump and the exposition-format file), it compares the metric value and the number of root spans for each series, prints one line mismatch<탭><handler><탭><code><탭><지표 값><탭><스팬 수> (the placeholders are a tab, the handler, a tab, the code, a tab, the metric value, a tab, and the span count) for each series that differs and ends with exit code 1, and if everything is the same it prints one line ok<탭><계열 수> (the placeholders are a tab and the number of series) and ends with 0. Save the output of running the checker on /opt/app/tracelab/tp_metrics/metrics/pay-api.prom in /root/tp-metrics/08-agree.txt.

The checker does not need otel, so write it to run under the system python3. A series only on the metrics side and a series only on the span side are both mismatches, so you must loop over the union of the two key sets. The grader also runs your checker on /opt/app/tracelab/tp_metrics/metrics/pay-api-broken.prom, which was deliberately written wrong, so you must not judge by file name or by particular values.