Where Distributed Tracing Breaks
The Structure of One traceparent Line
In one line
Distributed tracing boils down to passing the traceparent header on to the next call. Auto-instrumentation does it for you, but it breaks the moment your code creates a new request object.
Why this was needed
People often say "we installed Istio, so we have tracing," and that is wrong. A sidecar can create spans for the requests it sees, but the sidecar does not know that the request service A received and the request A sent to B belong to the same trace. Connecting them is the application's job.
The W3C standard header looks like this.
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
^^ ^------------ trace-id ---------^ ^-- span-id --^ ^^
버전 32자리(트레이스 전체) 16자리(이 구간) 플래그
The final 01 is the sampling flag. 1 means "record this trace" and 0 means "do not record it." This decision is made once at the very front and propagated onward — that way a trace never ends up with only half of it kept.
How it works
When service A calls B, there is exactly one thing it has to do.
- Read the
traceparentof the incoming request - Create a new span-id, leave the trace-id as it is, and attach it to the outgoing request
Auto-instrumentation does this for you. But it breaks in the following cases.
| Where it breaks | Why |
|---|---|
| A new HTTP client is created directly in code | A path the instrumentation library does not wrap |
| Handing off through a queue or message | There are no HTTP headers — you must put it in the message attributes yourself |
| A new thread or coroutine is started | The context is thread-local and does not follow |
| Batch processing | It is not one request : one trace |
Queues in particular are common. When you hand off through Kafka, you put traceparent in the message headers, and the consumer has to take it out and restore the context. If you do not, the producer-side trace and the consumer-side trace become strangers to each other.
Sampling
Recording everything is expensive. So we sample. There are two approaches.
Head sampling — decides "record this" at the very front. It is cheap and simple, but you cannot pick out only the slow requests. That is because at the start you do not know whether it will be slow.
Tail sampling — decides after the trace is finished. "Store only those with errors or over 1 second." It picks exactly what you want, but the collector has to gather the whole trace in memory before deciding, so it is expensive, and you must route spans of the same trace to the same collector.
The practical combination is usually this — cut down to 10–20% at the head, force 100% of errors, and apply tail sampling only to the important endpoints.
Common misconceptions
"The more spans, the better" — if you create a span per function, one trace becomes thousands of spans, you cannot read it in the UI, and storage cost explodes. Put spans at network boundaries and slow work.
Putting personal data in attributes. Span attributes are stored and searched as they are. If you put a user's email or a token in, it piles up in plain text in the tracing backend.
What to record on a span
A span with only a name and a time is of little help in an investigation. It needs attributes that let you tell what that span did, so that you can find what the slow ones have in common. However, if you put in just anything, storage cost and personal data problems come along with it, so you need a criterion.
What is worth including is generally things by which you can split the outcomes.
- The call target (service name, normalized path, database name)
- The result (status code, error kind)
- The scale (number of rows fetched, body size, batch size)
- The conditions (cache hit or miss, which path it took, retry count)
Scale and conditions are especially valuable. "This query is slow" gets to the cause far more slowly than "this query is slow only when it returns ten thousand rows," and that distinction is possible only if you have recorded the row count as an attribute.
Conversely, there are things you clearly should not include. Besides personal data and credentials, you should also be careful about values with a practically unlimited number of distinct kinds. The tracing backend builds an index so you can search by attributes, so if you put in the whole request body or a path that has not been normalized, the index explodes. It is the same story we told earlier about metric labels.
It also helps to settle how you record errors. A span has a separate status that expresses success or failure, so setting that correctly comes first. Without it, "show only failed traces" does not work. Leave the content of an exception as a span event, but decide whether to put in the whole stack trace together with the storage cost. Usually the error kind and a one-line message are enough, and the details can live in the logs. As we said earlier, if logs and traces are connected, you can just cross over to the logs.
What really matters in practice
Most of the real value comes from connecting traces and logs. If you put trace_id in every log line, then when you find a slow trace you can find all the logs that request left in one go. Without this connection, tracing stays a pretty picture.