TT Lab
Get started
Learn Learning paths Courses

OTCA — OpenTelemetry Certified Associate

Between a successful flush and successful receipt

Continue in TT Lab

In one line

Instrumentation is not a single call but a sequence of boundaries. You must confirm separately that an exception was recorded and that it was marked as an error, that a span was ended and that it was exported, and that a receiver accepted it and that it can be queried in storage.

Why this was needed

There is a short-running order validation job. The job catches errors and turns them into responses, calls force_flush at the end, and then logs whether it succeeded. The log printed True, but the observability screen has no request. If you conclude in this situation that "the observability screen is slow," you may miss the defect of not ending an open span or the fact that the transport was rejected.

The preliminary experiment for this lab used the real Python SDK 1.44.0 and the official OTLP/HTTP exporter. We made our own receiver respond with HTTP 200 and 400 and compared. The exporter's result differed, SUCCESS and FAILURE, but the provider's force_flush was True in both cases. The explanation below is based on this observation and does not generalize that every language and version gives the same return value.

How it works

An exception event and an error status are separate

When Python code catches and handles an exception, that exception may not propagate out of the with block. In this case, do not guess what the automatic context management will record; check separately the events and the status left on the current span. This lab handles an exception caught in a span created with start_span, so it does not mix with the behavior of automatic context management.

from opentelemetry.trace import Status, StatusCode

try:
    validate_order()
except ValueError as error:
    span.record_exception(error)
    span.set_status(Status(StatusCode.ERROR, "order validation failed"))

record_exception leaves an event to investigate. ERROR decides what status to classify this work as. In the preliminary experiment, a span that recorded only the event had the UNSET status, and only the span whose status was set separately was ERROR. This distinction is useful when it shows up in an event search but is missing from the error rate or a status filter.

Not every caught exception necessarily means the work failed. If you recovered through a normal fallback path, you must choose the status that fits the business meaning. In this assignment, it was stated explicitly that an order validation failure is classified as a final error, so ERROR is required. Memorizing how to call the SDK and judging which business outcome to apply that call to are different kinds of learning.

The order of span.end and force_flush

A batching processor collects finished spans and exports them. If it arbitrarily completed and exported the information of a span still in progress, the actual work time or the last event could be wrong. So even if you flush while leaving an open span, that span is not ended automatically.

span.end()
flush_result = provider.force_flush(timeout_millis=3000)

In the preliminary experiment, the first flush with an open span was True, but the number of requests received was 0. After the span was ended and flushed again, 1 arrived. The first return value was not evidence that the open work was completed. If you do not distinguish "finished because there was nothing to export" from "finished after sending the span you wanted," you miss the cause of data disappearing in short jobs.

force_flush is not a function to memorize as one you call unconditionally after every span. In ordinary services you make use of batch processing, and at short runs or at boundaries where a process may be interrupted, you design the wait and shutdown policy. This lab uses a direct flush to reproduce the boundary in a short time. It did not verify batch processing performance comparisons or high-load operational guidance.

Five different kinds of success

Observation point What that fact lets you say What you still cannot say
End observation after span.end The span has ended It was sent
Exporter call An export was attempted The receiver accepted it
Exporter SUCCESS That exporter treated it as a success Persistent storage, final queryability
The receiver's acceptance record The experimental receiver accepted the request It was forwarded to or stored in another backend
Lookup by ID in the target store That data is visible in that store All spans were preserved without loss

In the HTTP 400 experiment, the request body did reach the receiver. So if you count only the span list in received, it is 1. But the receiver rejected the request, so accepted_spans is 0 and the exporter result is FAILURE. If you define arrival of the body as success, you miss exactly this counterexample. In the lab, the delivered function is required to judge the receiver's acceptance, and it is not used as a guarantee of persistent storage.

This distinction is a way of thinking that also applies to logs, metrics, and message queues. First write down at which point you were looking when you called it a success: a function return, insertion into a local queue, a network send, the other side's acceptance, or final processing complete. Since the contract for return values differs by system, you must not transplant the scope of success just because the words are the same.

Where does service.name go?

Distinguish between writing service.name in a span's ordinary attributes and setting service.name on the Resource. The Resource describes the service that produced the signal. In the lab you call with two different service names, so if you hardcode the one string from the example in your code, it is right for one case and wrong for the other.

From the received protobuf, compare not only service.name but also the trace ID. This is so that you do not mistake another request with a matching name for the success of the current request. It is an exercise in narrowing the confirmation "something appeared on the screen" down to the confirmation "the request I just sent passed through this boundary." In a real service, you must also manage the storage and access permissions of the identifiers used for this.

What it looks like in the field

If you investigate an incident where observability data disappears right before a batch job ends, start with three questions. Did you end every span? Did the exporter have a chance to run at the shutdown boundary? What were the results of that attempt and of the receiving side? Do not stop at simply increasing the wait time; compare which boundary's data changes.

Another pitfall is test teardown. If you call provider.shutdown in the test's finally, data that had been missed can be sent late. If you count that data toward the student code's success, code that forgot to flush also passes. This grader copies the observation taken right after the student function and cleans up resources after that. The cleanup process is necessary, but it must not stand in for the correct answer.

What you will do in the next lab

The SDK and exporter are preinstalled in a dedicated environment. No external service or API key is needed. In eight steps you fix, in turn, the wiring, sampling, parent policy, exception, end, receive judgment, service identification, and the overall report. When you run each file with run, you can see the actual observation together with the failure conditions. Grading runs on a copy of the code and does not change the current file, and the preparation for earlier steps does not overwrite partial answers that already exist.

Official references: Python instrumentation, OTLP exporter.