OTCA — OpenTelemetry Certified Associate
The Four Places Propagation Breaks, and the Fix for Each
In one line
The only mechanism that builds the tree called a trace is context propagation. The caller puts the current trace ID and span ID into the W3C traceparent header, and the receiver reads it and makes it the parent of its own span. The places where this chain breaks are almost always the same — thread pools, message queues, background jobs, and intermediate layers that strip headers.
Why this was needed
You open a production trace, and the request passed through six services but there are only two SERVER spans. Adding manual spans in this state is pointless. The work for this week is to restore propagation in the other four places.
How it works
Anatomy of traceparent
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
-- -------------------------------- ---------------- --
| | | |
| | | +-- flags: 01 이면 샘플링됨, 00 이면 저장 안 됨
| | +-- parent-id: 16자리 16진수(8바이트). 호출 단계마다 바뀐다
| +-- trace-id: 32자리 16진수(16바이트). 요청 전체에서 동일하다
+-- version: 현재 00
The Korean text in this code block says, in order, that flags 01 means sampled and 00 means not stored, that parent-id is 16 hex digits (8 bytes) and changes at every call step, that trace-id is 32 hex digits (16 bytes) and is the same for the whole request, and that version is currently 00.
tracestate is a companion header that carries vendor-specific extra information, and baggage is a separate specification that carries user-defined key-value pairs across service boundaries. Baggage is not a span attribute, so it is not stored automatically.
Break 1 — thread pools and executors. Whether propagation happens depends on the runtime and the API you use. A plain ThreadPoolExecutor.submit without an instrumentation wrapper does not automatically pass the caller's context in this lab environment. In contrast, Python 3.12's asyncio Task and to_thread provide Context copying, and you must distinguish when the coroutine is created from when it is actually scheduled. At a boundary with no automatic support, capture the context on the caller side, activate it inside the worker, and restore the worker's previous state after execution. In the later context boundaries lab, you verify this directly with two orders and a reused worker.
Break 2 — message queues. A queue is both a process boundary and a time boundary. Headers do not flow automatically as they do with HTTP, so the producer must inject the context into the message headers and the consumer must extract it. There is one important design judgment here. A batch consumer must use a Link, not parent-child. If you process 100 messages at once, there are 100 parents, but a span can have only one parent. For a pipeline with long queue wait times, consider a Link even when there is only one message. If you connect with parent-child, the duration of a single trace grows by the wait time and becomes hard for the backend to handle.
Break 3 — background jobs and schedulers. A cron job or background task has no parent. A common mistake is to force-attach the request context. The request has already sent its response and finished, yet a 30-second child gets attached to that trace, contaminating the request latency statistics. The prescription is to start a new root trace and leave the originating request as a Link.
Break 4 — intermediate layers that strip headers. When a proxy, WAF, API gateway, or CDN filters headers with an allowlist approach, traceparent silently disappears. Nothing is left in the logs, and the symptom is "a new trace starts behind the gateway." Check that traceparent, tracestate, and baggage are in the allowlist, and if legacy systems using B3 headers are mixed in, configure multiple propagators.
The boundary between resource attributes and span attributes also appears on the exam.
| Category | Describes | Examples |
|---|---|---|
| Resource attribute | The entity that produced the telemetry | service.name, k8s.pod.name, host.name |
| Span attribute | That one unit of work | http.route, db.system, cart.item_count |
High cardinality in span attributes is not an explosion problem, unlike in metrics. Order IDs, user IDs, and query parameters are values that exist to be put into traces, and without them a trace is a picture you cannot filter. The problem lies in two places. One is when converting spans to metrics (an attribute you put in the spanmetrics dimensions becomes a metric label as is), and the other is the size of the span itself.
service.instance.id has high cardinality, but it is a resource attribute, so it is not a problem in traces and logs. However, if you promote it to a metric label, the time series are multiplied by the number of Pods, so it is common to remove it in the Collector for the metrics pipeline only.
What it looks like in the field
The cheapest way to check whether propagation is alive is to imitate a gateway. Send a request with a known trace ID put directly into the traceparent header, and then look it up by that ID in the backend to see whether the spans attach beneath it. If they do not, it is one of the four above.
Also keep a checklist of the criteria for saying that instrumentation is done. When you pick one production request and open its trace: does the number of services involved match the number of SERVER spans? Does the root span duration match the response time in the gateway access log? Is the largest self time under 20% of the total? Can logs be searched by trace ID? Are the top 20 span names route templates rather than IDs? Is work that crosses a queue connected as a single trace or via a Link? The last item — killing the Collector once and checking that the app's error rate and latency do not wobble — can be known only by actually doing it.
What to look for in the next check
This module is a conceptual module, so it has no lab. After the quiz confirms the traceparent fields and the criteria for choosing a Link, the last module moves on to sampling policies and buffer memory calculation.