OTCA — OpenTelemetry Certified Associate
Libraries Know Only the API; the Application Plugs In the SDK
In one line
An OpenTelemetry client is divided, for each signal, into four kinds of packages: API, SDK, semantic conventions, and Contrib. Instrumentation code depends only on the API, and the application owner installs the SDK and assembles providers, processors, and exporters. Following the specification overview and the client design principles, this article explains why this separation (composability) is needed, how the SDK pipeline is connected, and how an agent that does not modify code plugs in the same API and SDK.
Why this was needed
Instrumentation is, in the documentation's words, a "cross-cutting concern." Observability code gets mixed into web frameworks, DB clients, and message queue libraries. But the library author cannot know which backend the application using the library will send telemetry to, or whether it will send any at all. If a library drags in a specific implementation, even applications that do not want it become heavier.
So the design principles document requires three things. The API must be clearly separated from the implementation, third-party libraries must depend only on the API, and the final application developer must be able to decide how to configure the SDK — or whether to use it at all. To that end, it states that the API and the SDK MUST be provided as independent artifacts.
How it works
The API's minimal implementation
The API package is self-contained. The application must build and run even without an SDK, so the API contains a minimal implementation (no-op). The documentation stresses that the return values of this minimal implementation must be valid — createSpan() must not fail and must return a non-null Span, and the caller must not need to care whether the minimal implementation is currently running. It must also impose almost no performance burden. This is how the requirement that "an instrumented library can be used in an application that does not use OpenTelemetry" is met, and it removes the need for a framework to ship separate "instrumented" and "uninstrumented" editions.
The rule aimed at instrumentation authors is firm: "Instrumentation authors MUST NOT directly reference any SDK package of any kind. They reference only the API."
Inside the SDK — from provider to exporter
When an SDK is installed, it replaces the minimal implementation. The SDK is divided again into two parts — protocol-independent common logic (batching, attaching process information, and so on) and the protocol-bound exporter. Exporters have minimal functionality so that a vendor can easily attach its own protocol. The default exporters the specification requires of an SDK are OTLP (logs, metrics, traces), standard output, and in-memory (for testing), with Prometheus added for metrics and Zipkin for traces. Vendor-specific exporters are not put in the client.
The skeleton of the assembly has the same shape for every signal.
| Signal | Provider | What it creates | Unit recorded | Processor → exporter |
|---|---|---|---|---|
| Trace | TracerProvider | Tracer | Span | SpanProcessor → SpanExporter |
| Metric | MeterProvider | Meter → Instrument | Measurement | Aggregation state → MetricReader/Exporter |
| Log | LoggerProvider | Logger | LogRecord | LogRecordProcessor → LogRecordExporter |
According to the trace SDK specification, when you create a TracerProvider you configure SpanProcessors, an IdGenerator, SpanLimits, and a Sampler. The default sampler is ParentBased(root=AlwaysOn). The TracerProvider's Shutdown calls Shutdown on all registered processors, and ForceFlush propagates the same way.
The processor chain
A SpanProcessor is a hook into the span lifecycle — OnStart when a span starts, OnEnding just before it ends, OnEnd after it ends, and Shutdown and ForceFlush. They are called in the order registered, and OnEnd begins only after OnEnding has finished for all processors. There are two built-in processors.
- Simple processor: Passes a span to the exporter as soon as it ends.
- Batching processor: Collects finished spans in a queue and sends them in bundles. Its parameters are
maxQueueSize(default 2048; spans are dropped when it overflows),scheduledDelayMillis(default 5000),exportTimeoutMillis(default 30000), andmaxExportBatchSize(default 512; must be no larger thanmaxQueueSize). It exports when the queue reaches the batch size, when the delay has elapsed, or whenForceFlushis called, and the exporter'sExportcalls are serialized so they do not overlap.
Tracer ──► Span(끝) ──► [SpanProcessor 1] ──► [SpanProcessor 2: Batching] ──► SpanExporter(OTLP)
│ OnStart/OnEnding/OnEnd 훅
The log SDK has the same structure. You register a LogRecordProcessor with the LoggerProvider, and a Simple or Batching processor passes records to a LogRecordExporter (for example OTLP). The specification says the built-in processors handle "batching and transformation."
Agents — plugging in the same things without changing code
The zero-code instrumentation concept defines what an agent does as "adding the capabilities of the OpenTelemetry API and SDK to an application." The method differs by language — bytecode manipulation, monkey patching, eBPF. What gets instrumented is the libraries you use (requests and responses, DB calls, message queues), not your own code, and to instrument your own code you need code-based instrumentation. Configuration is done with environment variables and language-specific means, and the only thing you need to get started is a service name.
The Java agent is a single opentelemetry-javaagent.jar. Add -javaagent:path/to/opentelemetry-javaagent.jar to a JVM of Java 8 or later, and it injects bytecode dynamically and captures telemetry from many libraries. You can configure it with any of these: system properties such as -Dotel.service.name=..., environment variables such as OTEL_SERVICE_NAME and OTEL_TRACES_EXPORTER, JAVA_TOOL_OPTIONS, or a properties file specified with otel.javaagent.configuration-file.
Python uses monkey patching. pip install opentelemetry-distro opentelemetry-exporter-otlp fetches the API, the SDK, and two tools, and opentelemetry-bootstrap -a install scans site-packages and picks and installs the instrumentation libraries that match the installed packages (for example flask → opentelemetry-instrumentation-flask). You run it with opentelemetry-instrument python myapp.py and configure it with arguments such as --traces_exporter console,otlp or OTEL_* environment variables. The documentation states flatly that the distro package is required for auto-instrumentation to work.
What it looks like in the field
One team's in-house HTTP client library depended directly on the SDK. The batch jobs that used this library started an OTLP exporter even though they had nowhere to send telemetry, and they finished late because they waited on Shutdown at exit. When the library was changed to depend on the API, it became a no-op in batch jobs, and in web services the SDK that the application had assembled was used as is.
Another team had a problem where spans intermittently disappeared. When traffic spiked, the Batching processor's queue (default 2048) overflowed and dropped spans. They had to weigh increasing the queue against adjusting maxExportBatchSize and scheduledDelayMillis to raise the export rate.
What comes in the next article
The next article covers the seven kinds of metric instruments and temporality, the fields of a log record and the bridge API, and the schema URL. Then the quiz checks the reason for separating the API and SDK, the processor chain, the defaults of the Batching processor, and how agents work.