Where Distributed Tracing Breaks
Is a request that succeeded on the third try a success or a failure?
In one line
If you do not decide how many spans to draw for one retried request, and which span gets the error, the dump quietly lies.
Why this was needed
The person in charge of payments asked, "If it tried three times and succeeded on the third, is that a success or a failure?" The answer is both. To the user it is a success, and to the upstream service it is two failures. The problem was that the trace was drawn so that it could say only one of the two.
The code at that time wrapped the retry loop in a single span. On screen there was one charge span at 184 milliseconds, with a normal status. The fact that those 184 milliseconds contained one 60-millisecond timeout, one 20-millisecond 503, and two waits was recorded nowhere. All it could show someone asking why it got slow was "this span took a long time."
There was also a team that went to the other extreme. They created a span per attempt and put an error status on the failed attempts, and this time the error rate on the dashboard suddenly doubled. The failures users experienced had not increased; only the number went up. Nobody had told them in advance that if you count spans, retries get counted twice.
How it works
To sum up, there are four things to decide.
| Decision | Options | What this lab picks |
|---|---|---|
| Splitting spans | The whole retry as one / one per attempt | 1 wrapping span + N attempt spans |
| Where to put the error | The wrapping span / the failed attempt span | Only on the failed attempts |
| Status of an eventual success | Leave it / an explicit OK | Explicitly write that it is OK |
| Unit of counting | Span / logical request (the root) | Logical request |
The wrapping span is the one thing the user went through. If it was tried three times and finally succeeded, the user experienced a success, so this span is OK. An attempt span is one call that actually went out to the upstream. A failed call must be left as a failure so that you can see the upstream's health. Because the two layers answer different questions, their statuses go separately too.
Once this distinction is in place, where to count the error rate is decided on its own. If you use spans as the denominator, the more a request retried, the more the denominator and numerator grow together and failures inflate, while if you count only root spans, the failures users experienced come out as they are. It really happens that the same data yields 53.8% and 25.0% — in step 4 of the lab you produce those two numbers yourself.
The wait time needs a place to go too. The backoff between attempts is not in any attempt span, so it looks like only the leftover gap when you subtract the children's intervals from the wrapping span's interval. Do not leave people to guess what that gap was; it has to be written down as an event or an attribute. In step 5 of the lab you measure that gap yourself and match it against the record.
Attributes are useful only when nailed down as rules. The retry count and the last failure reason are properties of one logical request, so they go on the wrapping span, and which attempt it is differs per attempt, so it goes on the attempt span. The idempotency key goes on both with the same value — because only when you can see that it went out several times with the same key can you sort out duplicate-processing incidents. HTTP instrumentation already defines a standard attribute with the same meaning, http.request.resend_count, so before inventing a name yourself, it is better to look first at the HTTP span semantic conventions. The rules for the values of error.type, used for classifying failures, are also written in the error attribute registry.
Let us be clear about one thing not covered here. Which API records a caught exception on a span and how the status code is wired is covered by the SDK lifecycle module. This module's question comes before that — how many spans to express one task that was tried several times, and on which of them to put the error. How to set the status is in the Trace API specification.
We also write down what this lab environment cannot judge. The Pod has neither an OpenTelemetry Collector nor a tracing backend. So how tail sampling picks out failed attempts, and what shape retries take when folded on a backend screen, cannot be confirmed here. All we can see is the JSONL dump that writes out the spans the SDK exported, and all judgments are made from the structure and attributes of that dump. Elapsed times vary by a few milliseconds depending on the machine, so we look at them only as relationships, not absolute numbers.
What it looks like in the field
The sentence that comes up most often in incident retrospectives is "thanks to retries, the users felt nothing," and in many cases there is no way to check whether that was true. Without attempt spans you cannot count how often the upstream stumbled, and without wrapping spans you cannot count whether users actually experienced failures. Only when both layers exist can you say in numbers "the upstream was bad but the users were fine."
The opposite incident is common too. One team had put in three layers of retries — the client library three times, the service above it three times, and the gateway twice. When the upstream stumbled once, eighteen calls actually went out, yet the trace showed only one wrapping span. The moment they created attempt spans, those eighteen became visible, and that day they cut the retry layers down to one. The instrumentation exposed a design flaw.
What you will do in the next lab
With an upstream that deterministically fails twice and succeeds on the third try, you first see what the dump loses when retries are put in one span. Then you create a span per attempt and put the error only on the failed attempts, and from the same data calculate the span-based error rate and the request-based error rate to see how far apart they are. You record the wait time as events and match it against the gap, settle the attribute rules in a table, and then hook the same rules into a second service whose loop we cannot fix. Finally, you harden those rules into a linter and actually catch a dump that breaks them.