Where Distributed Tracing Breaks
Eliminate the causes of a missing span one by one
Goal
You compare the two pieces of evidence from a reported case to turn the missing requests into a number, narrow the scope by what they have in common, reproduce the four candidate causes yourself to confirm how differently each leaves its trace in the dump, organize that difference into a classification table and harden it into a diagnostic script, and see that running it on a second case with a different cause gives a different answer.
Why it matters
There are many reasons a span can be missing, but the symptom on screen is one — it is absent. So if you start fixing by guesswork, days go by as you raise the sampling ratio and then touch the exporter. We really did spend two days that way, and only then matched the logs and the dump by request identifier, and the fact that everything missing was a request on one path came out on the spot. There are two reasons diagnosis comes first. If you do not turn how many are missing into a number, you cannot tell whether things improved even after a fix, and if you do not look at what the missing ones have in common, you cannot choose between a problem where you must touch global settings and a problem where you must look at one path's code. The four candidates leave different traces in the dump, so if you make those traces yourself once, from then on a single dump tells them apart. Finding and fixing a defective wiring is the SDK lifecycle module's job, and what you build here are a classification table and a diagnostic script.
Steps
- The evidence for one case is in
/opt/app/tracelab/tp_missing/case1/— the request recordapp.logand the span dumpspans.jsonl. Match therequest_id=value of a log line against the span attributerequest.idin the dump, and find the requests that are in the log but not in the dump. Leave the result in two files. In/root/tp-missing/01-missing.txt, three lines — afterlogged=, the number of requests in the log, aftertraced=, the number of requests found in the dump, and aftermissing=, the number of missing requests. In/root/tp-missing/01-missing-ids.txt, write the identifiers of the missing requests in ascending order, one per line. - In the same case, look at where the missing requests are concentrated. Write four tab-separated columns in
/root/tp-missing/02-shape.tsv. First, one line per path,route<탭><경로><탭><로그 건수><탭><없는 건수>(the placeholders are a tab, the path, a tab, the log count, a tab, and the missing count), in ascending order of path name, then one line per minute,minute<탭><HH:MM><탭><로그 건수><탭><없는 건수>(the placeholders are a tab, the time, a tab, the log count, a tab, and the missing count), in ascending order of time. The last line isverdict<탭><route 또는 minute><탭><가장 많이 빠진 값><탭><그 값에서 빠진 건수>(the placeholders are a tab, route or minute, a tab, the value with the most missing, a tab, and the number missing at that value) — the line that chooses which axis they are concentrated on. - Make two causes that leave nothing in the dump yourself.
/root/tp-missing/sampling.pyuses the materialtracelab.tp_missing.samplers.drop_requests(["e-02", "e-05"])as the sampler and handles the six items ofwebapp.REQUESTS(default dump path/root/tp-missing/03-sampling.jsonl)./root/tp-missing/early_exit.pyhandles the same six without the sampler, but when it ise-05's turn it ends the process withos._exit(0)(default dump path/root/tp-missing/03-exit.jsonl). In both, the root span name isGET <경로>(the placeholder is the path), you pass the attributerequest.idwhen starting the span, and you callwebapp.work(tracer, req)inside. Then write two lines with three tab-separated columns in/root/tp-missing/03-nothing.tsv— the first line issamplingand the second isearly-exit, the second column is the identifiers of the requests missing from that dump joined with commas, and the third column istailif those missing ones are in a run from the end of the six, and otherwisescattered. - Create
/root/tp-missing/unfinished.py(default dump path/root/tp-missing/04-unfinished.jsonl). Handle the same six, but for the two itemse-02ande-05, create the root span withtracer.start_span(...)and do not end it (do not callend()). Those two must still create their child spans normally, so pass the context when calling, as inwebapp.work(tracer, req, context=trace.set_span_in_context(span)). The other four are done the same way as in step 3. After running it, the dump will have 10 span lines, and two of them will have aparent_idthat is not anyspan_idin the dump. - Create
/root/tp-missing/broken_parent.py(default dump path/root/tp-missing/05-split.jsonl). Handle all six normally, but for only the two itemse-03ande-06, call the child withwebapp.work(tracer, req, context=Context())to attach it to an empty context (from opentelemetry.context import Context). After running it, there will be 12 span lines and not a single missing request, yet two traces will appear that have not one span carryingrequest.id. - Looking at the four dumps you made in the previous three steps, write four lines with four tab-separated columns in
/root/tp-missing/06-fingerprints.tsv. The first column is the cause name, in ordersampled-out,unfinished,early-exit, andbroken-parent. The second column is whether the root span of the missing request is in the dump,noneorpresent. The third column is the shape of the child spans, one ofnone(there are none),orphan(there are some but the parent they point to is not in the dump), anddetached(there are some but they have become the root of a different trace). The fourth column is the distribution of missing requests, one ofscattered,tail, andnone(there are no missing requests at all). - Create
/root/tp-missing/classify.py. When run aspython3 classify.py <app.log> <spans.jsonl>(the placeholders are the log and the dump), it prints two lines —verdict=<원인 이름>andmissing=<없는 요청 수>(the placeholders are the cause name and the number of missing requests). Check the rules in this order. (1) If there is a span whoseparent_idis not anyspan_idin the dump,unfinished. (2) If there is atrace_idthat has not a single span with therequest.idattribute,broken-parent. (3) If there is not a single missing request,ok. (4) If the missing requests are in a run from the last line of the log,early-exit. (5) Otherwise,sampled-out. Save the output of running the script you built on/opt/app/tracelab/tp_missing/case1/as it is in/root/tp-missing/07-verdict.txt. - Run the same script on the second case
/opt/app/tracelab/tp_missing/case2/and save the output in/root/tp-missing/08-verdict.txt— a different answer from the first case must come out. Then leave an investigation record for the next person to read in/root/tp-missing/08-report.md. Put four headings in this order and write at least 60 characters under each —## 무엇이 없었나(what was missing: how many were missing in the two cases and where they were concentrated),## 어떻게 갈랐나(how you told them apart: by which traces you eliminated the candidates),## 원인(the cause: write the verdict names of the two cases as they are), and## 다음 사람에게(for the next person: what to do first if the same report comes again). Bothcase1andcase2must appear somewhere in the text.
Notes
- The working directory is
/root/tp-missing. If it does not exist, create it first. - Always run the instrumented programs as
/opt/otel-lab/bin/python <파일>(the placeholder is the file). The systempython3does not have the OpenTelemetry SDK. Conversely, run programs that only read logs and dumps with the systempython3. - The materials are the case folders
/opt/app/tracelab/tp_missing/case1to/opt/app/tracelab/tp_missing/case5(each withapp.logandspans.jsonl), the experiment request and child span helper/opt/app/tracelab/tp_missing/webapp.py, and the experiment sampler/opt/app/tracelab/tp_missing/samplers.py. The generator that made the cases is/opt/app/tracelab/tp_missing/make_cases.py, the shared wiring is/opt/app/tracelab/dump.py, and the dump-reading helper is/opt/lab/checks/_tplib.py. - Common mistake: running the program twice without deleting the dump file. The dump is appended to, so the spans double.
- Common mistake: attaching the attribute used for the sampling decision later with
set_attribute. The sampling decision is made when the span starts, so the sampler sees only the attributes passed at that time. - Sampling concepts · Trace SDK specification (ForceFlush, Shutdown, ShouldSample) · W3C Trace Context · Traces concepts · Python instrumentation docs
Compare the log and the dump to make a list of what is missing
The evidence for one case is in /opt/app/tracelab/tp_missing/case1/ — the request record app.log and the span dump spans.jsonl. Match the request_id= value of a log line against the span attribute request.id in the dump, and find the requests that are in the log but not in the dump. Leave the result in two files. In /root/tp-missing/01-missing.txt, three lines — after logged=, the number of requests in the log, after traced=, the number of requests found in the dump, and after missing=, the number of missing requests. In /root/tp-missing/01-missing-ids.txt, write the identifiers of the missing requests in ascending order, one per line.
The request identifier is attached only to server spans. Child spans (db.query) do not have it, so collect the request.id in attributes across the whole dump into a set. A log line is 열쇠=값 pairs separated by whitespace (the placeholders are the key and the value), so after split() you only need to look at the pieces that start with request_id=. Reading the dump does not need otel, so run it with the system python3.
Narrow the scope of the investigation by what the missing ones have in common
In the same case, look at where the missing requests are concentrated. Write four tab-separated columns in /root/tp-missing/02-shape.tsv. First, one line per path, route<탭><경로><탭><로그 건수><탭><없는 건수> (the placeholders are a tab, the path, a tab, the log count, a tab, and the missing count), in ascending order of path name, then one line per minute, minute<탭><HH:MM><탭><로그 건수><탭><없는 건수> (the placeholders are a tab, the time, a tab, the log count, a tab, and the missing count), in ascending order of time. The last line is verdict<탭><route 또는 minute><탭><가장 많이 빠진 값><탭><그 값에서 빠진 건수> (the placeholders are a tab, route or minute, a tab, the value with the most missing, a tab, and the number missing at that value) — the line that chooses which axis they are concentrated on.
Each log line contains route=, and for the time, five characters starting from the 11th character of 2026-09-16T09:00:00Z at the very front of the line are HH:MM. One axis will be entirely concentrated on one value and the other will be evenly spread — the concentrated one is the verdict. You already made the list of missing requests in step 1, so just use it.
Two causes that leave no trace at all are told apart only by distribution
Make two causes that leave nothing in the dump yourself. /root/tp-missing/sampling.py uses the material tracelab.tp_missing.samplers.drop_requests(["e-02", "e-05"]) as the sampler and handles the six items of webapp.REQUESTS (default dump path /root/tp-missing/03-sampling.jsonl). /root/tp-missing/early_exit.py handles the same six without the sampler, but when it is e-05's turn it ends the process with os._exit(0) (default dump path /root/tp-missing/03-exit.jsonl). In both, the root span name is GET <경로> (the placeholder is the path), you pass the attribute request.id when starting the span, and you call webapp.work(tracer, req) inside. Then write two lines with three tab-separated columns in /root/tp-missing/03-nothing.tsv — the first line is sampling and the second is early-exit, the second column is the identifiers of the requests missing from that dump joined with commas, and the third column is tail if those missing ones are in a run from the end of the six, and otherwise scattered.
The sampler decides when the span starts, so it cannot see a request.id attached later with set_attribute — pass it with start_as_current_span(이름, attributes={...}) (the placeholders are the name and the attributes). The module os must be imported in advance with import os to use os._exit. In both dumps, the missing requests will have not a single line for either root or child. Delete the file before making the dump again — the dump is appended to.
A span that was never ended leaves children with no parent
Create /root/tp-missing/unfinished.py (default dump path /root/tp-missing/04-unfinished.jsonl). Handle the same six, but for the two items e-02 and e-05, create the root span with tracer.start_span(...) and do not end it (do not call end()). Those two must still create their child spans normally, so pass the context when calling, as in webapp.work(tracer, req, context=trace.set_span_in_context(span)). The other four are done the same way as in step 3. After running it, the dump will have 10 span lines, and two of them will have a parent_id that is not any span_id in the dump.
A with statement calls end() for you when leaving the block, so to make a span that is never ended you must not use with. tracer.start_span only creates and does not even set it as the current span — so for a child to find its parent you must pass the context by hand. A span that is not ended is not handed to the exporter, so it does not appear in the dump at all.
If the parent context is broken, traces split with no missing requests
Create /root/tp-missing/broken_parent.py (default dump path /root/tp-missing/05-split.jsonl). Handle all six normally, but for only the two items e-03 and e-06, call the child with webapp.work(tracer, req, context=Context()) to attach it to an empty context (from opentelemetry.context import Context). After running it, there will be 12 span lines and not a single missing request, yet two traces will appear that have not one span carrying request.id.
An empty Context() has no current span, so a span started inside it cannot find a parent and becomes the root of a new trace. What makes this cause different from the previous three is that the comparison table catches nothing — so what you must count is not missing requests but traces without a request identifier. You can see how many traces there are with python3 /opt/lab/checks/_tplib.py summary <덤프> (the placeholder is the dump).
Harden the different trace of each cause into a classification table
Looking at the four dumps you made in the previous three steps, write four lines with four tab-separated columns in /root/tp-missing/06-fingerprints.tsv. The first column is the cause name, in order sampled-out, unfinished, early-exit, and broken-parent. The second column is whether the root span of the missing request is in the dump, none or present. The third column is the shape of the child spans, one of none (there are none), orphan (there are some but the parent they point to is not in the dump), and detached (there are some but they have become the root of a different trace). The fourth column is the distribution of missing requests, one of scattered, tail, and none (there are no missing requests at all).
For three of the four lines, the dumps you made yourself in steps 3, 4, and 5 show the answer as they are. The confusing ones are the first and third lines, which look identical from the dump alone and differ only in the fourth column. The grader rereads the dumps you made and checks that they match each line of the table — it is not a table to write from memory but a summary of your dumps.
Harden the same judgment into a script and run it on the case
Create /root/tp-missing/classify.py. When run as python3 classify.py <app.log> <spans.jsonl> (the placeholders are the log and the dump), it prints two lines — verdict=<원인 이름> and missing=<없는 요청 수> (the placeholders are the cause name and the number of missing requests). Check the rules in this order. (1) If there is a span whose parent_id is not any span_id in the dump, unfinished. (2) If there is a trace_id that has not a single span with the request.id attribute, broken-parent. (3) If there is not a single missing request, ok. (4) If the missing requests are in a run from the last line of the log, early-exit. (5) Otherwise, sampled-out. Save the output of running the script you built on /opt/app/tracelab/tp_missing/case1/ as it is in /root/tp-missing/07-verdict.txt.
The order matters — a span that was never ended also produces "traces without a request identifier," so if you do not check (1) before (2), the two causes get swapped. You count missing the same way for any cause (requests in the log whose request.id is not in the dump). The grader runs your script on other case folders under /opt/app/tracelab/tp_missing too, so you must not judge by file name or by a particular identifier.
Run it on a second case with a different cause and leave an investigation record
Run the same script on the second case /opt/app/tracelab/tp_missing/case2/ and save the output in /root/tp-missing/08-verdict.txt — a different answer from the first case must come out. Then leave an investigation record for the next person to read in /root/tp-missing/08-report.md. Put four headings in this order and write at least 60 characters under each — ## 무엇이 없었나 (what was missing: how many were missing in the two cases and where they were concentrated), ## 어떻게 갈랐나 (how you told them apart: by which traces you eliminated the candidates), ## 원인 (the cause: write the verdict names of the two cases as they are), and ## 다음 사람에게 (for the next person: what to do first if the same report comes again). Both case1 and case2 must appear somewhere in the text.
The dumps of the two cases look alike on the surface — both are missing 16. What differs is where in the log those 16 are, and whether children with no parent remain in the dump. When you write the record, do not write only the conclusion but what you looked at and what you ruled out. What the next person needs is not the answer but the order.