Where Distributed Tracing Breaks
When someone reports that a trace is missing
In one line
There are four reasons a span can be missing but only one symptom, so do not fix by guesswork; compare the request record with the dump and tell the causes apart by the different traces each one leaves.
Why this was needed
The report always comes in the same sentence. "I looked for the trace by this order number and it doesn't show up." We once spent two days starting from that one sentence. The first two days we suspected the sampling configuration. We raised the ratio and waited a day, and the reports kept coming in just the same. Next we suspected the exporter and increased the batch size, and that wasn't it either. Only on the third day, after matching the logs and the dump by request identifier, did we find that everything that was missing was a request on one path.
Looking back, it was something we could have done on the first day. There were two things we did not do. First, we did not turn "how many are missing" into a number. Because we moved on only the words "it doesn't show," we could not tell whether things had improved even after a fix, and so we had the same suspicion twice. Second, we did not look at what the missing ones had in common. When everything is missing and when only one path is missing, the causes are completely different, yet we touched the global settings without making that distinction.
How it works
Diagnosis is four steps. Compare to make a list, narrow by what they have in common, eliminate candidates, and confirm what is left.
The first step is the list of what is missing. The application log has one line per request, and the server spans in the span dump carry the same request identifier as an attribute. If you take the difference of the two sets, you get "requests that are in the log but not in the dump." Only with this number does everything afterward become real.
The second step is narrowing. Count the missing requests by path and by time period. If they are concentrated on one path, look at that path's handling code; if they are concentrated in a time period, look at what happened at that time; and if they are evenly scattered, look at the global settings. This one table cuts the investigation scope to a tenth.
The third step is elimination. There are four candidates, and the traces the four leave in the dump differ from one another.
| Candidate cause | Root span | Child spans | Distribution of missing requests |
|---|---|---|---|
| It was left out of the sample | Absent | Absent | Scattered |
| The span was never ended | Absent | Present and pointing to a parent, but that parent is not in the dump | Scattered (usually one path) |
| The process ended first | Absent | Absent | In a run from the end of the log |
| The parent context was broken | Present | Present, but has become the root of a different trace | There are no missing requests |
Sampling and early termination both leave no trace at all, so you cannot tell them apart from the dump alone. What separates them is the distribution. Sampling scatters across the whole range, while if the process dies, everything after that time is missing whole. A span that was never ended, on the contrary, leaves a very distinct trace — the child was exported, but the parent it points to is nowhere in the dump. A broken parent context is an entirely different symptom. There is not a single missing request, yet a separate trace floats around that has not a single span carrying the request identifier. On screen it looks like "the trace has only one span."
The fourth step is hardening the same judgment into a script. If a person draws the table every time, the next report costs two more days. If you write the rules in code, in a fixed order, the next person gets the same answer with one command. The reason to put the rules in an order is that the traces can overlap — a span that was never ended also produces "traces without a request identifier," so it must be checked before the parent context judgment.
What exactly sampling throws away and why children follow the parent's decision are laid out in the sampling concept docs, and why exporting before the process ends is an explicit call is in the ForceFlush and Shutdown sections of the Trace SDK specification. The original for how parent context comes in and goes out is W3C Trace Context.
Let us be clear about what is not covered here. Finding and fixing a defective SDK wiring is the SDK lifecycle module's job. What this module teaches comes before that — the procedure for telling from the data which of several causes it is, and the output is not fixed code but a classification table and a diagnostic script.
We also write down what this environment cannot judge. The lab Pod has neither an OpenTelemetry Collector nor a tracing backend. So how that trace is drawn on a backend screen, and what the collector drops along the way, cannot be confirmed here. What we look at is only two things, the JSONL dump that writes out the spans the SDK exported as they are, and the logs the service left, and all judgments are made by comparing the two.
What it looks like in the field
The one seen most often is the third. In batch jobs and short command-style programs, the last few requests are always missing, and the report comes in as "it occasionally drops." When you compare with the logs, you see right away that what is missing is always at the end. The one fact that it is not scattered immediately rules out the sampling candidate.
The second most frequent is the fourth. If you do not pass the context inside a queue worker or a callback, the spans in it become the root of a new trace. Nothing is missing, so the comparison table catches nothing, yet people report that "the trace is half." What you have to count then is not missing requests but traces without a request identifier.
What you will do in the next lab
You start with two pieces of evidence (the request log and the span dump) from one reported case. First you compare them to count the missing requests, and split them by path and by time period to see where they are concentrated. Then you reproduce the four causes yourself to confirm the trace each leaves in the dump, and organize the differences into a classification table. You move the rules you organized into a diagnostic script and run it on the case, and finally run the same script on a second case with a different cause to confirm that a different answer comes out, and leave an investigation record.