OTCA — OpenTelemetry Certified Associate
The same function, but another order’s tag
In one line
A context problem is not only whether a value exists, but when it was copied and where it was restored. If you do not treat a Task, a deferred coroutine, and an ordinary thread pool as the same execution boundary, you can narrow down why requests get mixed up by observing just a few lines of code.
Why this was needed
There is a service that handles two orders at the same time. When it receives order A, it sets request=order-a and schedules a stock check. Soon afterward, the caller's value is changed to another request. Yet the stock check log prints the later request's tag instead of A's. The function arguments are fine and there is only one process, so at first it looks like a sorting error in the logging server.
This lab reproduces that situation without an external server. In an environment with Python 3.12 and OpenTelemetry SDK 1.44.0 installed, it runs real asyncio and a real ThreadPoolExecutor. It is not a method of writing the correct sentence into a fake log file. It reads and compares the baggage at the moment the student function ran the callback and the caller's baggage after the function returned.
If the earlier module asked "was the span received?", here we ask "whose context was that work run in?" Even if a lot of observability data arrives, if different requests are wrongly connected, the analysis of an incident's cause will be wrong. That is why you must confirm the meaning of the connection before increasing the collection volume.
How it works
Creating a value versus attaching it as the current value
OpenTelemetry's baggage.set_baggage returns a Context containing the value. It is not the same operation as attach, which applies that value to the current execution. If you detach with the token that attach returned, you go back to the context from before you entered. Read the contracts of the Python Context API and the Baggage API as distinct. Do not assume that a callback reads a value just because you created a Context.
The first step of the lab is a small scope that shows the callback a value while preserving the original caller. The caller's value is caller, and the inner value is order-a. You run separately the case where the callback returns normally and the case where it throws a ValueError. If the inner value is right but the outside is left as order-a, the code is only half right. Discarding the return value or replacing the exception with a new one also breaks the contract of the business code.
If you put the restore only at the very end of the success path, an exception in the middle skips that line. So in the lab, the original value must be observed on both the normal path and the exception path. A test that runs only one request and then exits the process can hardly find this defect. You need a test that also looks at what runs after the value is left behind.
Write down three points in time
The lab's preliminary experiment first attached alpha and created the work. Next it changed the caller to beta and then waited for that work. Even though the execution order looks the same, the value the worker read differed depending on which API was used to create the work.
| What was done at alpha | Measured result when waiting after switching to beta |
|---|---|
| create_task(coro) | The Task read alpha |
| to_thread(function), only creating the coroutine | The thread read beta |
| create_task(to_thread(function)) | The thread read alpha |
This table is the result of this module's Python 3.12 experiment. All three rows are "work that runs later," but the moment at which the context is captured differs. The important point for the second row is that only a coroutine object was created and its body had not yet run. The third row scheduled that coroutine as a Task in the current context, separating it from the caller's change.
The Python documentation explains that create_task by default uses a copy of the current Context, and that merely calling a coroutine does not schedule it to run. It also explains cancellation cleanup in connection with try/finally. Python 3.12 asyncio. If you check this contract against the actual values above, you can explain the problem more accurately than with slogans like "async makes it happen automatically" or "threads always lose it."
The lab's start_task and start_thread are functions that preserve the request at the time of scheduling. It is not the correct answer to permanently revert the caller's current value so that the worker sees the right one. The caller must continue to be later-caller. You can tell propagation from leakage only by observing both sides together. Return the scheduled Task so that the caller can await and cancel it. Do not extend this exercise to a background-task pattern that discards both the result and the reference.
An ordinary executor is a separate experiment
The ordinary ThreadPoolExecutor.submit in the preliminary experiment did not pass alpha along automatically. When the submitting side copied with contextvars.copy_context and ran the function inside the run of that copy, alpha was observed. Then, when an ordinary function was submitted to the same worker, there was no value. We checked separately that "it was passed once" and that "it did not remain for the next work."
If you copy inside the worker function, you copy the context after it has already crossed the boundary. Student code fails if the execution order is wrong, even if it contains the name copy_context. The grader does not search for the word; it compares actual return values. The task of this step is to make a separate copy for each submission and preserve the state of a reused worker. Creating a new thread every time can hide the leakage problem on reuse, so two orders are checked with the same worker.
Cancellation is also the end of a request
When a customer drops the connection or a parent task is canceled, the callback may not return normally. In step 5, two requests are gathered at the same point with an Event and then allowed to continue, and the values in the overlapping interval are checked. Next, exceptions and cancellation are triggered and the restore is confirmed. At the end, cancel is called on a Task that is actually waiting. This is so that we do not claim to have verified the external cancellation path just because the case of raising CancelledError directly passed.
If you catch a cancellation and return it as an ordinary success, the inner value is restored but the caller does not learn that it was canceled. The contract of this lab is to restore the value while passing the exception and the cancellation on to the caller. Cleanup and hiding an error are not the same thing.
What it looks like in the field
The following is the investigation order you will practice in this lab. First split the request in two and use different synthetic identifiers. Next record the values just before the call, right after scheduling, inside the callback, and after the return. If the normal path is right, add exceptions and cancellation. If there is a thread pool, reuse the same worker. Also leave the execution environment and the API names in the result so that the next person does not mistake it for the behavior of a different runtime.
This test needs no real personal data or production tokens. With only order-a and order-b you can tell the propagation timing from the leakage. If you use real customer identifiers when you only need the values to differ, the investigation material itself becomes something new to manage.
What you will do in the next lab
In steps 1–5, you fix in turn the small-scope restore, Task scheduling, thread scheduling, executor submission, and the cancellation cleanup of overlapping requests. Look at the observations and checks of each run and explain which boundary was wrong. The next reading separates the fact that a value was delivered exactly from the fact that the value can be trusted. A successful delivery is not a successful authentication.