Where Distributed Tracing Breaks
Deciding What to Throw Away
In one line
Sampling is not a device for cutting cost but a design for deciding what you will see. If you get it wrong, the trace you actually need is not there.
Why this was needed
A trace creates dozens of spans per request. At 1,000 requests per second, that is billions of spans a day. If you store them all, storage cost exceeds the cost of the application.
But if you blindly keep only 1%, then when an outage happens there is a 99% chance the trace of that request does not exist. So sampling is approached not as "how much should we cut" but as "what must we be sure to keep."
Three approaches
| Approach | When it is decided | Advantage | Cost |
|---|---|---|---|
| Head | At request start | Cheap and simple, easy to propagate | Cannot pick out only slow requests |
| Tail | After the trace completes | Picks exactly the errors and slow ones | Expensive, because spans must be gathered |
| Rate limiting | An upper limit of N per second | Cost does not spike even in a surge | Fewer samples in the surge period |
The practical combination is usually this.
- Cut the baseline to 10–20% at the head
- Force 100% of errors (turn on the sample flag of
traceparentand propagate it) - Catch the slow ones with tail sampling only on the important endpoints
The sampling decision must be propagated
This is where people get it wrong most often. Sampling must be decided per trace, but if each service decides separately, the trace falls into pieces.
프런트(20%) → API(20%) → 결제(20%)
각자 정하면 세 서비스가 모두 남길 확률 = 0.2³ = 0.8%
→ 완전한 트레이스는 거의 안 남고, 조각난 스팬만 쌓인다
The last byte of the W3C traceparent header is the sample flag. You must decide once
at the very front, and afterward only follow that decision.
traceparent: 00-4bf92f...36-00f067aa0ba902b7-01
↑ 01 = 샘플됨, 00 = 아님
The OpenTelemetry SDK's ParentBased(root=TraceIdRatioBased(0.2)) implements this.
If there is a parent, it follows the parent's decision, and only when there is none (when it is the root) does it roll a 20% die.
If you use only TraceIdRatioBased, each service rolls on its own and the 0.8% problem above appears.
This is also why TraceIdRatioBased decides not by chance but by a hash of the trace ID.
The same trace ID gives the same answer no matter which service computes it, so if the configuration
is the same, the decisions agree on their own.
How to keep 100% of errors
"We keep all the errors" is easy to say but impossible with head sampling. When a request starts, you do not know whether it will fail. There are two workarounds.
The application turns on the flag. At the moment it detects an error, it forcibly marks the current span as a recording target. It cannot save the parent spans that have already passed, but everything below is kept.
Use tail sampling. The collector decides after seeing the whole trace, so it is accurate. A policy looks like this.
tail_sampling:
decision_wait: 10s
policies:
- name: errors # 오류는 전부
type: status_code
status_code: {status_codes: [ERROR]}
- name: slow # 2초 넘는 것은 전부
type: latency
latency: {threshold_ms: 2000}
- name: baseline # 나머지는 5%
type: probabilistic
probabilistic: {sampling_percentage: 5}
Policies are evaluated with OR — if any one matches, it is kept.
The hidden cost of tail sampling
The collector gathers the whole trace in memory and decides afterward. So two things follow.
- Memory: it accumulates as much as wait time (usually 5–30 seconds) × traces per second
- Routing: spans of the same trace must go to the same collector. So you put a load balancer keyed on
trace_idin front of the collectors. If you do not, traces are fragmented and the decisions are wrong
If you turn on tail sampling without knowing these two things, the collector dies of OOM, and the traces in flight while it is down vanish entirely.
Common misconceptions
"Raising the sampling ratio makes it more accurate" — metrics must be independent of sampling. Measure request counts, error rates, and latency with metrics, and traces are the tool for seeing "where did that request spend its time." If you compute ratios from traces, the sampling bias goes straight in.
"Keeping only errors is enough" — slow requests that are not errors give more information. Errors also stay in the logs, but a request that got a normal response and took 3 seconds cannot be explained without a trace.
What really matters in practice
The sampling decision must be made once at the very front and propagated to the end. If a middle service decides again on its own, only half of the trace remains.
The last flag of traceparent (01/00) carries that decision. Auto-instrumentation respects this, but if code that builds its own HTTP client writes the header anew, the decision is reset at that point.