The Question Metrics Cannot Answer
Summary
The fact that each service is individually fast and the fact that a single request that passed through them in order was slow do not contradict each other.
Why this was needed
A report comes in that the payment screen takes 2 seconds. The gateway p99 has risen to 1.8 seconds. If you open the p99 of the ten services behind it one by one, all are normal.
In this situation, metrics cannot give an answer in principle. This is because metrics are aggregates and cannot reconstruct the path of an individual request. There is no guarantee that service A's p99 and service B's p99 belong to the same request. What you need is not per-service statistics but the whole path of one request.
How it works
A trace is a tree of spans that share one trace ID. Each span has a name, start and end times, a parent span ID, attributes, and a kind. The kind matters in practice — SERVER is the receiving side and CLIENT is the calling side. The time difference between the CLIENT span and the SERVER span is the time spent in the network and queuing, and you can know that value only when both are present.
The only mechanism that builds this tree is context propagation. The standard header is the W3C traceparent, and its format is as follows.
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
^버전 ^트레이스 ID(32hex) ^부모 스팬 ID(16hex) ^샘플 플래그
The trace ID is the same throughout the request, and the span ID changes at every hop. Even just following these two rules lets you join logs by trace ID, and that alone solves half the problem.
Tracing answers three questions. First, where one request spent its time. If out of millions of requests a day, only 1% spend 1,612ms on a coupon lookup, an average dashboard does not budge. Second, the problem of repeated calls. Even if an individual query is fast at 4ms, if it runs 340 times in one request, it is 1.4 seconds. This shows up only if you count spans. Third, conditional paths — cases where it becomes slow only under a particular combination of a specific tenant, a specific flag, and a specific cache miss.
What you meet in the field
The points where propagation breaks are almost fixed. Your own HTTP client that automatic instrumentation could not wrap, code that hands work to a thread pool or worker, message queues, and third-party proxies that strip headers. If a trace is strangely short, look at these four places first.
The concept of self time is also worth knowing. It is the value obtained by subtracting the time of the child spans from a span's total duration. If this is large, it means there is in-process work that automatic instrumentation cannot see — serialization, sorting, compression — hiding there.
Even before you finish attaching tracing, at least put in a correlation ID. Create an ID at the request entry point, put that field in every log, and send it in the headers of every downstream call. This alone dramatically reduces incident investigation time.
How to decide sampling
If you store the traces of every request, the cost is unaffordable. So you keep only some, and how you choose almost decides the value of tracing.
The simplest method is to roll a die at a set rate when a request comes in. It is easy to implement and the cost is predictable, but it has a decisive weakness. Slow requests and failed requests are dropped with the same probability. Yet those are exactly the requests you want to investigate. With 1% sampling, only 1 out of 100 problematic requests remains.
So a method was devised that decides whether to keep a request after it ends. It gathers a trace for a short time until it is complete, and keeps it if it was slow or had an error, and discards it otherwise. You can keep exactly what you want, but there is a cost. You must remember the whole trace, so the collector side needs memory and state, and when that collector becomes several instances with scale, you must arrange the routing so that spans of the same trace go to the same collector.
In practice, people usually mix the two. The default is a low rate to get the picture of normal times, and errors or slow requests are always kept. And the decision is made once at the very front of the request and the result is propagated in a header. If each service rolls its own die, traces break from the middle and leave fragments that are useless anywhere. The last flag of traceparent does exactly this.
One more thing to decide is the retention period. Traces are larger in volume than logs, so it is hard to keep them long. Instead, metrics extracted from traces can be kept for a long time, so a workable combination is to keep the raw data for only a few days and extract per-service and per-path latency distributions as metrics for long-term storage. For questions that compare with several months ago, metrics answer, and for why this request was slow now, traces answer.
What you will do in the next lab
You start a 3-tier chain of services and create and propagate a traceparent yourself. You validate the format, confirm that the span ID changes at every hop, join the logs of the three services by trace ID, find the point where propagation broke, and finally calculate the slowest interval and the self time.