TT Lab
Get started
Learn Learning paths Courses

OTCA — OpenTelemetry Certified Associate

The Spans Survive Even When the Chain Breaks

Continue in TT Lab

In one line

Even when propagation breaks, each service's spans remain intact. So the symptom shows up not as "tracing doesn't work" but as "the back part of the request can't be seen."

Why real requests must flow

The earlier modules covered propagation and the Collector. But the places where those labs ran had no requests flowing, so the phenomenon of propagation breaking never happened at all. Even though this is the area OTCA asks about most.

A break is silent

Even when propagation breaks, each service's spans remain intact. When you open the screen, the traces are there, and the names and times are normal. They just don't connect into one.

So the symptom is not "tracing doesn't work" but "why can't I see the back part of this request?" The cause is much harder to find.

1. 라이브러리가 헤더를 안 붙인다      직접 만든 HTTP 클라이언트가 흔하다
2. 프록시가 헤더를 지운다             허용 목록에 traceparent 가 없다
3. 큐를 건너갈 때 잃는다              메시지 본문에 넣어 손으로 이어야 한다
4. 스레드·비동기 경계에서 끊긴다      컨텍스트가 따라가지 않는다

The Korean text in this code block lists four causes, in order: the library does not attach the header (a hand-written HTTP client is common), a proxy strips the header (traceparent is not in the allowlist), the context is lost when crossing a queue (it has to be put in the message body and connected by hand), and the chain breaks at thread and async boundaries (the context does not follow along).

The last two are especially silent. With HTTP you can tell by looking at the headers, but with queues and async there is no obvious place to look.

To notice, you need to keep the ratio of spans with no parent as a metric. If a service always creates a root span, the chain is broken right before it.

When the span arrived but is not on the screen

This is a spot where we actually got bitten while building the lab. The Collector log said spans had come in, but nothing was on the screen.

The timestamp was the cause. When we put in a small number, it was indexed as 1970 and could not be found in the default query range (recent). And Jaeger's query API returns an empty result if you do not give it a time range.

Connecting propagation across a queue by hand

With HTTP, the instrumentation library attaches the headers, but with a message queue nobody does it for you. You have to put the context into the message when publishing and take it out when consuming.

from opentelemetry import propagate, trace

# 발행 쪽 — 현재 컨텍스트를 헤더 맵에 주입한다
def publish(body):
    carrier = {}
    propagate.inject(carrier)                  # traceparent 가 여기 들어간다
    queue.send({"body": body, "otel": carrier})

# 소비 쪽 — 꺼낸 컨텍스트를 부모로 삼는다
def consume(msg):
    ctx = propagate.extract(msg.get("otel") or {})
    with tracer.start_as_current_span("handle", context=ctx,
                                      kind=trace.SpanKind.CONSUMER):
        handle(msg["body"])

If you do not pass context= to start_as_current_span, a new root span is created and the trace breaks at that point. This is the most common cause of a growing number of spans with no parent.

At an async boundary, also set the kind. If you mark PRODUCER and CONSUMER, the UI draws the queue segment differently, and the time spent waiting in the queue becomes visible.

Cardinality sets the cost

What you put into span attributes determines storage cost and query speed.

OK to put in Not OK to put in
http.route (/users/{id}) The full http.target (/users/48213)
db.system, db.operation The full SQL (with parameters)
messaging.destination The message body
Tenant ID (hundreds of values) User ID (millions of values)

If you put the path in as is, attribute values grow by the number of requests and the index explodes. Use a route normalized to a template. And never put personal data in attributes under any circumstances — traces usually have a loose retention policy and are viewed by many people.

What to instrument first

If you try to instrument everything, you can never start. There is an order.

  1. Entry points — HTTP servers and queue consumers. Even with just these, you can answer "which request is slow?"
  2. Outgoing calls — HTTP clients, DBs, caches. With these, you get "where is it slow?"
  3. Inside the application — pick only the heavy computation sections. If you create a span per function, only the cost grows and the screen becomes hard to read.

Items 1 and 2 are mostly obtained with auto-instrumentation. What you write by hand starts at item 3.

What really matters in practice

Keep the ratio of spans with no parent as a metric. It is practically the only way a person notices that propagation has broken. If a service always creates a root span, the chain is broken right before it, and then you can pinpoint the spot.

Queue and async boundaries must be connected by hand. With HTTP, the instrumentation library attaches the headers, but putting the context into the message body and taking it out is something nobody does for you. If there is a section that uses a queue, the propagation for that section must be built in at the design stage.

If spans don't show up, look at the timestamp first. If you put in a value that is not in nanoseconds, it is indexed as 1970 and can never be found in the default query range. If the Collector log says "received" but the screen is empty, it is almost always this case.

In the next lab, you verify these directly on a real Collector with real requests.