OTCA — OpenTelemetry Certified Associate
Auto-Instrumentation Gives You One Thing: the Network Boundary
In one line
Turning on auto-instrumentation gets you exactly one thing: the network boundary. Spans appear for HTTP servers and clients, gRPC, DB drivers, Redis, and message queue clients. What happens inside the process shows up as no span at all, and appears only as a gap in the parent span's self time.
Why this was needed
Teams that fail at instrumentation fail in almost the same way. They start by planting manual spans in the code, and two weeks later they have 300 spans but the traces are still cut off at service boundaries. The order that works is fixed.
| Order | What to do | If you skip it |
|---|---|---|
| 1 | Turn on auto-instrumentation and confirm data arrives | All later debugging becomes guesswork |
| 2 | Fix the resource attributes | Changing them later breaks continuity with past data |
| 3 | Verify propagation across service boundaries | Adding more spans still leaves traces fragmented |
| 4 | Add manual spans only where self time is large | The blanks in auto-instrumentation stay forever |
| 5 | Move processing and sampling to the Collector | Every service must be redeployed whenever the policy changes |
The key point is that step 3 comes before step 4. Adding manual spans while propagation is broken just cuts already fragmented traces into even smaller pieces.
How it works
With auto-instrumentation on, you get a trace like this.
SERVER checkout-api POST /v1/orders 1421ms
├─ CLIENT GET http://auth.internal/verify 31ms
├─ CLIENT SELECT carts WHERE id = ? 6ms
├─ CLIENT redis GET promo:rules:t-8871 2ms
├─ CLIENT POST http://payment.internal/charge 74ms
└─ (나머지 1308ms 는 어떤 스팬에도 속하지 않음)
The last line says it all. Auto-instrumentation tells you where the problem is not through a 1,308ms gap. That gap is the self time, and manual spans go only there.
What auto-instrumentation can never see is the following: CPU work inside the process (serialization, compression, template rendering, encryption), lock waits and connection pool waits, GIL contention and event loop delay, calls to third-party SDKs that have no instrumentation package, and branches in business logic.
There are five places to put manual spans.
- Loop and batch boundaries — record the iteration count as an attribute
- Cache lookups — if you record whether it was a hit as an attribute, cache efficiency is immediately visible in the trace
- Calls to third-party SDKs that have no instrumentation package
- Waits on locks, queues, and connection pools
- CPU-heavy sections — serialization, compression, report generation
After adding them, the earlier gap is filled like this.
├─ INTERNAL checkout.apply_promotions 1298ms
│ ├─ INTERNAL promotion.load_rules cache.hit=false 1241ms <-- 여기
│ └─ INTERNAL promotion.evaluate evaluated=812 54ms
You only need to follow one span naming rule. Names must be low cardinality. Use GET /v1/orders/:id, not GET /v1/orders/A-99183, and all concrete values go into attributes. The backend groups by span name to build latency statistics and the service graph, so if IDs appear in names, that whole aggregated view collapses.
Resource attributes cannot be fixed later. Unlike span attributes, resource attributes are attached to every signal the process emits, and the moment you change service.name, the connections to dashboards, alerts, the service graph, and past data are all cut.
| Attribute | Example | Can it be changed? |
|---|---|---|
service.name |
checkout-api | Practically no |
service.namespace |
commerce | Hard |
service.version |
2.7.1 | Changes with every deployment |
deployment.environment.name |
prod | No |
service.instance.id |
Pod name | Changes with every restart |
Two things people often get wrong here. First, the name of the environment attribute is deployment.environment.name. The old name deployment.environment is no longer used, and if the names differ you get two separate attributes, and a dashboard variable reads only one of them. Second, service.name is per service, not per deployment unit. The moment you call a canary checkout-api-canary, a ghost node appears in the service graph.
We recommend starting with parentbased_always_on as the sampler. If you turn on ratio sampling from the start, you cannot tell whether a missing trace is an instrumentation problem or a sampling effect. Even when you move to ratio sampling, always use parentbased_traceidratio. If each service makes its own independent probability decision instead of following the parent's decision, traces get cut off midway.
What it looks like in the field
Instrumentation breaks apps in a handful of recurring ways.
| Symptom | Cause | Response |
|---|---|---|
| Memory keeps growing after deployment | The exporter queue fills up because the backend is slow | Set an explicit queue limit and switch to going through a Collector |
| Only some spans arrive | The process dies without a flush on shutdown | Call shutdown and extend the termination grace period |
| A single span is hundreds of KB | The entire request body was put into an attribute | Set an attribute value length limit |
| Ghost node in the service graph | The canary was deployed under a separate service.name | Keep service.name fixed per service |
The SDK has no default limit on attribute value length. The number of attributes is limited to 128 by default, but the value length is unlimited, so if a whole request body goes in, transmission and storage costs follow directly. It is safer to set an explicit limit.
What you will do in the next lab
Under /root/otca-sdk/, you write an SDK environment variable file, set the resource attributes according to the conventions, and create and actually apply a Deployment that injects the Pod name and namespace with the Kubernetes Downward API. At the end, you tidy up a list of span names and write a linter yourself that catches names with IDs embedded in them.