OTCA — OpenTelemetry Certified Associate
Tail Sampling Is Not Free — Three Costs
In one line
Head sampling decides at a moment when nothing has happened yet, so it discards rare events exactly as rarely as they occur. Tail sampling decides after seeing the outcome, but it incurs three costs — the network is not reduced, it eats a lot of memory, and it requires trace ID-based routing. What was discarded at the head cannot be recovered at the tail.
Why this was needed
Storing everything does not make economic sense for most organizations. The problem is what you throw away. Suppose you investigate 5 errors per hour with 1% head sampling.
보존되는 오류 트레이스 = 5 × 0.01 = 시간당 0.05건
1건을 보려면 평균 20시간 대기
The Korean text in this code block says, in order, that the retained error traces equal 5 × 0.01 = 0.05 per hour, and that seeing one takes an average wait of 20 hours.
This is the practical outcome of head sampling. When you need to investigate, that trace is not there, and the traces that are there are all normal requests with no reason to look at them.
How it works
Tail sampling gathers spans in a buffer until the trace is complete and then decides after seeing the outcome. It keeps the trace if there is an error, keeps it if it is slow, and keeps only a small amount of the rest. Each policy is a "condition to keep," and a trace is retained if any one of the policies matches. So to exclude health checks, you use invert_match, which flips the match.
| Item | Head sampling | Tail sampling |
|---|---|---|
| When the decision is made | When the root span is created | After decision_wait has elapsed |
| Retaining errors and slow requests | Only probabilistically | 100% by rule |
| Agent and network cost | Reduced by the ratio | No reduction |
| Backend storage cost | Reduced | Reduced |
| Collector memory | Negligible | A buffer of traces per second × wait time |
| Operational requirement | None | Trace ID-based routing is mandatory |
Cost 1 — the network is not reduced. To make a decision, every span must arrive at the Collector, so the only savings are in backend storage and indexing. The Collector's own CPU and network actually increase.
Cost 2 — memory. You must hold all spans during decision_wait. The formula is simple.
초당 트레이스 수 × decision_wait(초) × 트레이스당 스팬 수 × 스팬 크기 = 버퍼 메모리
예: 10,000 trace/s × 30s × 12 스팬 × 1.2KB
= 4,320,000 KB = 약 4,219 MB (여유 2배면 약 8,438 MB)
The Korean text in this code block says, in order, that buffer memory equals traces per second × decision_wait (seconds) × spans per trace × span size, and that the example works out to about 4,219 MB, or about 8,438 MB with a factor of 2 headroom.
num_traces is the upper bound on the number of traces kept in memory at the same time, and if it is exceeded, the oldest trace is forcibly decided. The value you need comes from the same multiplication — traces per second × decision_wait. With 10,000 trace/s and a 30-second wait, you need 300,000.
Cost 3 — routing. All spans of the same trace must arrive at the same Collector instance. If you scale the Collector out to several instances and put an ordinary load balancer in front, the spans of one trace are scattered across instances, each decides by looking at its own fragment, and the result is randomly truncated traces. The solution is to put a loadbalancing exporter layer in front that uses routing_key: traceID.
For the same reason, tail sampling is impossible in a DaemonSet. In a structure where a Collector runs on every node, the spans of one trace arrive scattered across several nodes, so each agent sees only part of the trace. Use tail sampling only in the Deployment (gateway) layer.
What it looks like in the field
The most expensive lesson was a case where traffic grew while num_traces was left at 50,000. The peak was 10,000 trace/s and decision_wait was 30 seconds, so the required value was 300,000. The limit was one sixth of that, so old traces were constantly forced to a decision, and with memory pressure piling on, the Collector kept restarting from OOM and lost five minutes of data. The metrics reported almost nothing wrong — the data was being sent "normally," and all that looked off was that only certain traces were missing from the backend.
There are two lessons. First, do the multiplication before you turn on tail_sampling. Second, if data disappears while the metrics are quiet, take tail_sampling out of the pipeline for a moment and see whether the problem reproduces. This is the fastest way to tell.
The practical combination is usually this. A service with very heavy traffic is reduced in advance at the head to 10–50%, and tail sampling is layered on top to catch errors and slow requests. What has already been discarded at the head cannot be recovered at the tail, so the principle is to set the head ratio as high as you can afford.
What you will do in the next lab
Under /root/otca-sampling/, you write five tail sampling policies: outcome-based policies that keep 100% of errors and latency, a policy that uses invert_match to exclude health checks, a composite policy that joins a VIP tenant and latency with and, and a probabilistic policy for the rest. Then you configure the front-end loadbalancing exporter, calculate the buffer memory yourself and leave it in a file, and finally write an OpenTelemetryCollector CR that holds those values as they are.