Where Distributed Tracing Breaks
Can you reproduce the incident from this span alone?
In one line
A good span is one that lets you reproduce the same failure from that one line alone. The attributes are that reproduction input, the events are what happened and when in between, and the links are which other trace asked for this work.
Why this was needed
Payment failures clustered at dawn. When we opened a trace, the span had http.route=/checkout, http.response.status_code=502, and 1842 milliseconds. Auto-instrumentation had done all of that. But there was nothing we could do with this span. How many items were in the cart, whether the cache was empty, how backed up the queue was, how many times payment had been retried — not a single value needed for reproduction was there.
Conversely, one team went with "let's put everything in for now" and caused an incident. They put the whole request body into attributes, and customer emails and auth tokens piled up as they were in the observability backend, which every developer could see. After that the team made what goes into attributes a code review item.
This lab deals with what lies between those two. It is not about how to configure the SDK, nor in what order to decide service.name, but what to fill the attribute slots with.
How it works
First, flip the question. Instead of "what should we put in," ask "what else would we need to reproduce the same failure holding only this span." Write down the answer, and that is your attribute list. Reproduction input (item count, body size), the state at that time (cache hit, queue depth), and what we did (retry count) usually go here.
Next, separate values you must not include. Values such as contact details, card numbers, and auth tokens are not kept as the original. But throwing them out entirely causes trouble in an investigation, so you keep them after converting them in one of three ways.
| Original | How to keep it instead | Questions you can still answer |
|---|---|---|
| Email or account | Hash (first few characters) | "Does it repeat for the same user?" |
| Card number | Category (brand) | "Does it happen only with a particular card brand?" |
| Auth token | Length | "Did the token arrive truncated?" |
Third, a fact that has a point in time is an event, not an attribute. That a retry happened twice is a single number, so it is an attribute, but when and why the first retry happened is a record with a timestamp, so it is an event. If you keep the same fact only as an attribute, the order disappears, and if you keep it only as an event, aggregation becomes hard. The two are not competitors; they have different roles — things you count go in attributes, and the moment something happened goes in an event.
Fourth, some relationships are not parent–child. When a job stacked in a queue is processed later, that job was created in another trace, and if you attach the processing span as a child of that trace, you get a strange parent that ends hours later. This is when you use a link. A link says only "this span is related to that span" and leaves the parent slot empty. Many teams write the relationship as an attribute string (parent_trace_id=...), but then the tools cannot connect it and a person has to find it by eye. Here we only try out once that "a relationship is a link, not an attribute," and how to design traces across a queue is covered separately in a later module.
The last is cost. If you count how many distinct values each attribute key has, its nature splits. An order number is normally different for every request (an identifier), and a card brand should normally be three kinds (a category). The problem is when the original leaks into a slot that should be a category — if an order number is embedded in an exception message, one key suddenly takes thousands of values. So for each key, write down separately "keys with unlimited distinct values" and "categorical keys that should have few distinct values," and count the dump to find where they disagree.
If you leave what you have settled so far only as sentences, the next person will not read it. Make it a convention file with one line per key giving its name, type, allowed values, and the nature of its distinct values, and put alongside it a small checker that finds spans that break the convention. Then the convention follows along on its own when you instrument a new handler.
Let us also write down what cannot be judged in this Pod. There is no collector. In real operations the collector's redaction processor filters once more, but here that stage cannot be run, so you see only the result of the application filtering on its own. How attributes are indexed in the backend and how expensive that is also cannot be known in this Pod — instead we substitute the count of distinct values from the dump. And the SDK settings that truncate the number and length of attributes are outside this lab's scope.
What it looks like in the field
One service's investigations stopped at "we can't reproduce it" every time there was an outage. The spans did not even have a request identifier, so you could not find in the logs which request had failed. After adding six attributes (order number, item count, body size, cache hit, queue depth, and retry count), they could pick one failed span and put the same input in again, and investigation time dropped from hours to minutes.
Another incident was the exact opposite. Because of the practice of putting exception messages into attributes as they were, one key came to have tens of thousands of values, and the screen that searched by that key died of a timeout every time. The fix was not to delete the messages but to split them in two — short kinds that can be classified (declined and timeout) stayed as attributes, and the long sentences meant for people to read moved to the span status message and to events. Search became fast again and the sentences for people to read were kept as they were.
What you will do in the next lab
You receive one line of a failed request's span and first write down what is needed for reproduction but missing. You add those values as attributes, and convert values that must not be included into hashes, categories, or lengths. You move facts with a point in time, like retries and cache misses, into events, run two hundred requests, and count the distinct values per attribute key to make a budget table. Based on the leaking keys that show up there, you write a team convention file, and build a program that checks that convention and run it on your dump and on someone else's dump. Finally, while instrumenting a refund handler triggered by a queue according to the convention, you connect the trace that created that job with a link.