TT Lab
Get started
Learn Learning paths Courses

Where Distributed Tracing Breaks

Finding the spike you saw on the dashboard back in a trace

Continue in TT Lab

In one line

Only if metrics and traces have attributes with the same name and the same value, and the sampling ratio is written on the span, can you find again in a trace the peak you saw on the dashboard.

Why this was needed

The request that came out of an incident retrospective was simple. "The error rate panel spikes at 09:40, so let's open just one failed request from then." But we could not open it. The metrics had the label handler="/order/:id" and the spans had http.target="/order/8821". The names were different and the values were different. Nothing connected the two except people's heads, so there was no road from the peak to a trace.

Something worse came next. Someone said "can't we just count by spans?" and computed the error rate from the dump and brought back a number, 29%. The metrics pointed to 7.5% for the same period. It was a long time before we found the cause — that service was keeping every request that errored and only one in five of the successful ones. Since the sample was tilted toward errors, the ratio counted from that sample could not possibly be right.

How it works

To pair the two signals, three things are needed.

First, a common attribute must exist on both sides with the same name and the same value. Metrics cannot use the address as it is because of cardinality, so they use a registered route template. Traces may hold the original address as it is, but the attribute used for pairing must have the same value as the metrics. So on the span you keep two things together — http.route for pairing and the original address for investigation. The names are aligned as the same http.route by the HTTP span semantic conventions and the HTTP metrics semantic conventions, so it is better to use them as they are rather than invent your own.

Second, you must not count ratios from spans. Metrics count every request but traces keep only a sample. Even if the sample was drawn evenly, the count by spans is a fraction of the real one, and if the sample is tilted toward errors, the ratio itself is wrong wholesale. In the lab you produce 7.5% and 29.0% from the same data yourself.

Third, you must write the sampling probability on the span to be able to undo it. How many requests one span represents is the reciprocal of the probability that span was kept. A span kept with probability 1/5 represents five requests, and a span kept with probability 1 represents one. If you add up those weights, the counts and ratios largely come back — the standard calls this calculation the adjusted count, and it is defined in the probability sampling specification.

You must also clearly know what cannot be undone. A combination that never made it into the sample stays 0 — this is how a rare error vanishes from the dashboard entirely. Quantiles do not come back either. Weights fix counts but not the shape of a distribution, and the tail in particular swings badly the smaller the sample. So it is right to read counts and ratios from metrics and use traces to find one example.

The bridge from metrics to traces is choosing that example in advance. If you keep one representative trace_id per interval, you can go from the panel straight to that trace. The standard name for this approach is the exemplar, and the exposition format specification is in the Prometheus exposition format docs and the OpenTelemetry metrics data model. However, in this lab environment you cannot actually store or query exemplars. This image has no Prometheus at all, and the Prometheus in the observability lab Pod also runs with exemplar storage turned off. So here we go only as far as choosing the representatives and building that table, and we handle metrics as exposition-format text files.

We also write down the limits as they are. There is only one representative per interval, so you cannot jump to an ordinary request, and a request left out of the sample can never become a representative however slow it was.

Let us also be clear about what is not covered here. Choosing cumulative versus delta and configuring views in the metrics SDK is the metrics temporality module's job, and deciding which events to take as an SLI is the SLI module's job. What this module does lies between the two — instrumenting so the two signals can be paired, and its output is neither a metrics configuration nor an SLI definition but an attribute rules file and a checker that compares the two signals.

What it looks like in the field

The most common incident is the route values disagreeing. On the metrics side the framework puts in the route template automatically, while on the tracing side, instrumented by hand, the address is often put in as it is. Then a metric with nine series faces more than thirty bundles of spans, and there is no way to pair them. The dashboard looks fine, so nobody knows about this mismatch until an outage.

The second is a sampling policy that keeps all errors. It is put in with good intentions, and someone who counts ratios from that sample is sure to appear. If you write the sampling probability on the span, there is at least a way back, and if you do not, that data can never be used for ratio calculations.

What you will do in the next lab

First you count request counts, error counts, and latency distribution from the span dump alone. Then you compare with the exposition-format file the metrics side produced and make a table of how the names and values disagree, and fix the instrumentation so the two signals attach by the same key. Next you swap samplers to produce, from the same data, different error rates yourself, correct them back with the reciprocal of the sampling probability, and write down what even the correction cannot fix. Finally you build a bridge that picks a representative trace per interval, harden the rules into a file, apply them to a second service, and confirm with a checker that the two signals match.