OTCA — OpenTelemetry Certified Associate
OTLP Is a Protocol, Not a Store
In one line
OpenTelemetry is a project that standardizes how telemetry is created and sent. It does not store or display it. OTLP is the transport protocol for that, and you choose the backend, whether Tempo, Jaeger, or ClickHouse. Draw this boundary first, and the misconception "if I install OpenTelemetry, will I get a dashboard?" goes away.
Why this was needed
In the days when every vendor had its own instrumentation library, changing the backend meant re-instrumenting every service, because calls to a specific vendor's SDK were embedded in the code. OpenTelemetry cut that dependency out. The application speaks only to the OTel API, and where the data goes is decided in the SDK configuration or the Collector configuration. So replacing a backend becomes a configuration change rather than a code review.
How it works
Four signals each answer a different question.
| Signal | Question it answers | Typical limit |
|---|---|---|
| Traces | Where did this request spend its time? | Cannot show since when or what percentage is affected |
| Metrics | Since when, and how much has it gotten worse? | Cannot reconstruct individual requests |
| Logs | Why did it happen inside that? | Correlation across services must be wired up by hand |
| Profiles | Which code spent that time? | The most recently standardized signal |
Traces answer "where," metrics answer "since when and how much," and logs answer "why." You need all three for an investigation to reach the end. Copy a trace ID from a trace, paste it into a log search, and see whether that request's logs appear — this is the cheapest test of whether your instrumentation is connected.
Think of the components in three layers.
- API: The interface that application code calls. Without an SDK it is a no-op that does nothing. This is why a library author can add instrumentation depending only on the API.
- SDK: The implementation of the API. It has a sampler, processors (batching), exporters, and a resource.
- Collector: A relay process outside the app. It handles collection, processing, and routing. It is optional, but most setups include one.
OTLP is the protocol that flows between these. It has two transports, and the trap beginners fall into most often is here.
| Transport | Default port | OTEL_EXPORTER_OTLP_PROTOCOL |
|---|---|---|
| gRPC | 4317 | grpc |
| HTTP/protobuf | 4318 | http/protobuf |
If the port and protocol do not match, the connection itself fails, and the error appears only in the application log. The Collector-side metrics do not move from 0, so it is easy to misdiagnose this as "the Collector is not receiving." When no data arrives at all, this combination is the first thing to check.
There are rules for specifying the endpoint too. OTEL_EXPORTER_OTLP_ENDPOINT is the base address shared by all signals, and when you use HTTP the SDK appends a path such as /v1/traces. To use a different destination per signal, use a per-signal variable such as OTEL_EXPORTER_OTLP_TRACES_ENDPOINT; in that case you must write the full address including the path.
What it looks like in the field
The fact that the Collector comes in two distributions also trips people up in practice. Core contains only the essential components and is small, while Contrib includes a large number of community components and is much bigger. It is common to paste a component name seen in documentation or a blog post and have the Collector die with "unknown type," and most people first suspect a typo in the configuration. Far more often the real cause is a different distribution. You need the habit of first checking, with otelcol components, what the current binary contains and what stability level each component has.
Also, Jaeger v2 was redesigned on top of the OpenTelemetry Collector and receives OTLP natively. The Zipkin exporter is deprecated, so there is no reason to adopt it in a new project.
What to decide when you actually add instrumentation
OpenTelemetry is a broad subject, so after learning it you get stuck on where to start when you try to apply it. There is an order.
Turn on auto-instrumentation first. Most languages have an agent or library that creates spans for HTTP servers, clients, and DB drivers without any code change. That alone can answer "which request is slow?", and that is 80% of observability.
Then add spans by hand. What auto-instrumentation does not know is our business logic. Put spans only around meaningful units (order validation, stock checks, settlement calculations), and do not add one per function. The more spans there are, the harder they are to read and the more they cost.
Follow the conventions for attribute names. If you use the established names such as http.request.method, db.system, and service.name, tools build the screens for you. Names you invent make sense only in your own dashboards. Distinguish in-house attributes by adding a prefix.
Check propagation. The most common problem is a trace breaking as it crosses a service boundary. You need to check whether the traceparent header passes through proxies, queues, and batches. With message queues you often have to write the code that inserts and extracts the header yourself.
Decide on sampling from the start. If you send everything, the cost is unaffordable, and if you reduce randomly, the slow requests you care about do not survive. Tail sampling (choosing slow or erroring requests to keep after the request has finished) is the answer, but it uses memory on the Collector side. A reasonable start is "keep all errors and slow requests, and 1% of the rest."
Put a Collector in between. If the application sends directly to the backend, you have to redeploy everything whenever you change the backend. With a Collector, changing the destination, filtering, and retries all finish in that one place.
What to look for in the next check
This module is a conceptual module, so it has no lab. After the quiz confirms the signal boundaries and the OTLP protocol, in the next module you will write SDK environment variables and resource attributes yourself and inject an instance identifier with the Kubernetes Downward API.