Where Distributed Tracing Breaks
What to keep and what to drop
Goal
If you store every trace the cost becomes unaffordable, and if you cut at random, the requests you actually want to see disappear. You settle between the two with numbers.
Materials
/opt/lab/trace/spans.jsonl 요청 1만 건 (지속 시간·상태·서비스)
mkdir -p /root/trace && cp /opt/lab/trace/* /root/trace/ && cd /root/trace
head -2 spans.jsonl
What to keep
01-shape.txt 분포와 오류 건수
02-head.txt 무작위 1% 가 남기는 것
03-tail.txt 오류·느린 것을 전부 남길 때
04-cost.txt 저장 비용
05-memory.txt 꼬리 표본의 숨은 비용
06-decide.md 우리가 쓸 규칙
07-notes.md 왜 그런지
Look at the distribution first
Compute the total count, p50, p95, p99, and error count of spans.jsonl, save them in 01-shape.txt, and write one line describing the shape of this distribution.
python3 - <<'PY'
import json
d = sorted(json.loads(l)['duration_ms'] for l in open('spans.jsonl'))
n = len(d)
print(n, d[n//2], d[int(n*0.95)], d[int(n*0.99)])
PY
The mean tells you nothing in this distribution. Most are fast and a few are very slow, so the mean becomes a value in between that nobody experiences.
What does random 1% miss
Pick 1% at random from the whole set and save in 02-head.txt how many requests over p99 and how many errors remain in that sample. Write what a random sample misses.
import random; random.seed(7)
sample = random.sample(rows, len(rows)//100)
Around 100 remain, and of those, the ones over p99 will be one or two.
What you need in an incident investigation is exactly those one or two. A random sample keeps the common things well and loses the rare ones — but what we want to see is always on the rare side.
Choose after it ends
Save in 03-tail.txt how many you keep when you choose all the errors, all those over p99, and 1% of the rest, and write how many times more that is compared with random 1%.
It is called tail sampling because it decides after the request ends.
Add 129 errors + those over p99 (that are not errors) + 1% of the rest. You get around 300 — about three times random 1%, but you did not lose a single thing you wanted to see.
The sample ratio is the cost
Taking one trace as 4KB, at 1,000 requests per second and 30 days of retention, calculate the storage for (1) storing everything, (2) random 1%, and (3) tail sampling, and save it in 04-cost.txt.
It is requests per second × 86,400 × retention days × size per item × sample ratio.
If you store everything, a big number comes out even for a single day. If you start with "let's just turn everything on" without doing this calculation first, you find out on next month's bill.
The hidden cost of tail sampling
To do tail sampling, estimate what the collector has to hold and until when, and the memory needed for that, and save it in 05-memory.txt.
You can tell whether a request is slow or an error only after it ends. So the collector has to hold all of that request's spans in memory until it ends.
Memory needed ≈ requests per second × average duration (seconds) × size per item.
If this number is large, you must split the collector across several machines, and then spans of the same trace must go to the same collector for the decision to work. That is the point where running tail sampling gets hard.
Write it down as a rule
Write in 06-decide.md the sampling rule to use for our service. Decide how to treat errors, the threshold for slow ones, and the rate for the rest in numbers, and write the reason for each.
"Sample appropriately" is not a rule. It is a rule only if the next person can carry it straight into the configuration.
Also decide whether to set the threshold at p99 or at a fixed value (for example 1 second). p99 rises along with the service as it slows down, so while it is slowing down, nothing may get caught.
For the next person who reads this
Pick four or more of the things you saw here and summarize them in 07-notes.md. Write not what you did but why it is so.
Imagine it will be read by yourself when you next touch the sampling settings. "We chose tail sampling" does not help, but "random keeps the common things well and loses the rare ones, and what we want to see is always on the rare side" does.