OTCA — OpenTelemetry Certified Associate
The trace ID exists, but where is the span?
In one line
That a span has an identifier, that its content is recorded, and that it was handed to an exporter are different facts. If you confirm the three separately, you can narrow the problem "I added instrumentation code but can't see anything" using observation instead of guesswork.
Why this was needed
You added tracing to an order API. The developer found the trace ID in the log, and the code has start_span too. But the observability screen shows nothing. Restarting the Collector at this point means picking one of several possible causes without evidence. In reality, the processor may not have been connected to the provider, the sampler may have dropped the span, or the span may never have been ended. There are also cases where it was delivered to the exporter and then the receiving server rejected it.
This module covers the SDK side of those boundaries. Unlike the earlier lab where you hand-built OTLP JSON, here spans created by the official Python SDK are sent through a real processor and exporter. You look at the information visible in front of the exporter together with the actual HTTP receive result. The fact that the program ended without an error is not enough to judge that the data was delivered.
How it works
API, SDK, provider, processor, exporter
The application's instrumentation code creates spans through the API. The SDK's TracerProvider owns the runtime settings such as the sampler, the resource, and the processors. Getting a tracer does not automatically complete the delivery path. You must also decide which processor hands finished spans to which exporter. In the first lab, this one connection was deliberately left empty.
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
provider.add_span_processor(SimpleSpanProcessor(exporter))
tracer = provider.get_tracer("orders.instrumentation")
Here the name of the tracer is a name that distinguishes the scope of the instrumentation library. It is not the same place as service.name, which sets the name of the service. You can put the same string in both values by coincidence, but their roles do not become the same. If you keep the values separate from the start, you can read the origin even when several libraries instrument one service.
Three combinations of recording and sampling
The following are the results you will observe with this lab's official SDK and the default export processor. Processor observation means the call after the span has ended.
| Decision | Is it recording | sampled bit | Reaches processor | Reaches exporter |
|---|---|---|---|---|
| DROP | No | Off | No | No |
| RECORD_ONLY | Yes | Off | Yes | No |
| RECORD_AND_SAMPLE | Yes | On | Yes | Yes |
RECORD_ONLY is especially important. The name "records" makes it easy to read it as also being sent, but in the actual experiment there was only the processor's end observation, and the exporter's span list was empty. The purpose of observing information in memory and the purpose of deciding how much data to send remotely can be separated. Do not stop at memorizing this table; change one decision in the student code and run it to see which list grows and shrinks.
Even when the content is discarded by sampling, a valid span context can exist. So the trace ID in a log is a clue for searching, not a guarantee that a stored span exists. Conversely, the fact that is_recording became false after ending a span does not mean it was DROP from the start. You can tell the two cases apart only if you write down the moment at which you observed the state.
The root of ParentBased is not a switch for the whole trace
The policy you will build in the lab is one where a request with no parent is not recorded, and a request with a parent follows that parent's sampled decision. ParentBased(root=ALWAYS_OFF) expresses this policy. Even with root turned off, a child that comes from a sampled remote parent is recorded. If you interpret the root option as "a setting that turns off all spans," you end up diagnosing perfectly normal children as missing.
Conversely, ParentBased(root=ALWAYS_ON) does not automatically turn on the child of an unsampled parent. The root sampler is the decision for the case with no parent. There is a separate branch for each combination of remote or local parent and sampled or unsampled, and the defaults follow the parent's decision. This assignment runs both remote cases and both local cases, and it does not accept an implementation that happened to be right for one header as the correct answer for the whole policy.
If the parent is off, can the child never be on?
No. In the preliminary experiment, ALWAYS_ON and an explicitly changed remote-parent policy produced a sampled child with the same trace ID as an unsampled parent. Following the parent's decision is a sampler policy, and the identifier itself does not forbid the child from being recorded. Python's official sampling API also distinguishes always_on from parentbased_always_on.
However, this does not mean the content of a parent that was already discarded has been recovered. The attributes, events, and time spans that the parent did not record are still absent. You must distinguish that one child span has become observable from the trace of the whole request having become complete. In the field, expecting that "we forced sampling on, so now we'll see everything" starts a different kind of misdiagnosis.
What it looks like in the field
Imagine a service that normally collects only part of its root traffic receiving a request from an external partner. Whether the partner already sent a sampled context or not cannot be explained by looking only at the same root ratio setting. You check, in order, whether the context received in the request is valid, whether it was interpreted as remote, and whether the selected sampler follows the parent's decision.
To find out whether the cause is SDK sampling, comparing the counts before and after the SDK processor can be faster than just turning up the Collector logs. If both the start and end observations are missing, look at the creation path and sampling. If the end observation exists but there is no export, look at sampled and the processor connection. If the export is already confirmed, move to the transport boundary from there. Moving the place you observe one step at a time reduces unrelated infrastructure changes.
The lab's data is synthetic orders, not real customer information. Even in production, you must not unconditionally put raw orders, authentication headers, or personal data into span attributes for error investigation. First think about whether the minimum necessary identifiers and states are enough to tell the processing paths apart.
What to check next
The next reading distinguishes exception events and span status, span end and flush, and the boundary between receiving and storing. The lab that follows has you edit real Python functions to match the explanation. The grading criterion is not a string that makes a green light but the actual SDK's observation results.
Official references: Tracing SDK, Python sampling API.