OTCA — OpenTelemetry Certified Associate
Tail Sampling Policy and the Routing Layer
Goal
Design five tail sampling policies, put a trace ID routing layer in front of them, calculate the buffer memory yourself, and reflect that value in an Operator CR.
Why it matters
Tail sampling is not a feature where "turn it on and cost goes down." To make a decision, every span must arrive, so the network from the agent to the Collector is not reduced at all, and what shrinks is only backend storage and indexing. In exchange, you pay three new things: the memory to hold every span during decision_wait, the routing layer that gathers the same trace on one instance, and the complexity of operating that layer. So the decision to adopt it is made by multiplication, not by feel. Multiply traces per second by the wait time to get num_traces, and multiply that by the span count and size to get the memory. If you skip this calculation, you run with too small a limit, old traces are forced to a decision, and you meet the situation where data disappears while the metrics report nothing wrong.
Steps
- Create
/root/otca-sampling/tailsampling.yamland, inprocessors.tail_sampling, writedecision_wait: 30s,num_traces: 300000, andexpected_new_traces_per_sec: 10000. - Put two policies in
policies:name: keep-errors(type: status_code, withERRORinstatus_code.status_codes) andname: keep-slow(type: latency,latency.threshold_ms: 800). - Add a policy
name: drop-healthchecks. Usetype: string_attribute,string_attribute.key: http.route,valuesof/healthzand/readyz, andinvert_match: true. - Add a policy
name: vip-slow. It istype: and, andand.and_sub_policyhas two sub-policies — one withtype: string_attributematchingtenant.tierequal toenterprise, and one withtype: latencyandthreshold_ms: 300. Finally, addname: baseline(type: probabilistic,probabilistic.sampling_percentage: 2) to make a total of 5 policies. - In
/root/otca-sampling/loadbalancing.yaml, write the front-end routing layer. Inexporters.loadbalancing, setrouting_key: traceID, aresolver.dns.hostnamethat is a service address containingotel-tailsampler, andresolver.dns.port: 4317. The exporter ofservice.pipelines.tracesis just[loadbalancing], and you must not puttail_samplingin the processors. - In
/root/otca-sampling/buffer-memory.txt, write the calculation results on three lines:num_traces_required=(10,000 trace/s × 30s),buffer_mb=(10,000 × 30 × 12 spans × 1.2KB converted to MB and rounded), andbuffer_mb_with_headroom=(twice that). Each line is in the키=값form (key=value) with no spaces. - Create the namespace
otca-samplingand write the CR in/root/otca-sampling/collector-cr.yaml. UseapiVersion: opentelemetry.io/v1beta1,kind: OpenTelemetryCollector,metadata.name: otel-tailsampler,metadata.namespace: otca-sampling,spec.mode: deployment, andspec.replicas: 3. Insidespec.config, put thetail_samplingfrom steps 1–4 (the five policies as they are), and writespec.config.service.pipelines.traces.processorsso thattail_samplingcomes beforebatchandbatchis last.
Notes
- The calculation for step 6: 10000 × 30 × 12 × 1.2 = 4,320,000 KB. Divide this by 1024 and round.
- The structure of
spec.configin step 7 is identical to an ordinary Collector configuration file. Putreceivers,processors,exporters, andserviceinside it as they are. - Common mistake 1: leaving out
invert_match. Then only the health checks are kept. - Common mistake 2: putting
tail_samplingin the front-end routing layer as well. It would decide before the spans have gathered. - Common mistake 3: setting the CR's
spec.modetodaemonset. Spans are scattered across nodes and a decision is impossible.
tail_sampling defaults
Create /root/otca-sampling/tailsampling.yaml and, in processors.tail_sampling, write decision_wait: 30s, num_traces: 300000, and expected_new_traces_per_sec: 10000.
decision_wait is the time margin for all spans of a trace to arrive. If it is too short, you decide on an incomplete trace, and if it is too long, memory spikes. num_traces comes from the product of traces per second and the wait time.
Two outcome-based policies
Put two policies in policies: name: keep-errors (type: status_code, with ERROR in status_code.status_codes) and name: keep-slow (type: latency, latency.threshold_ms: 800).
These two policies are the reason tail sampling exists. They retain 100% by rule what head sampling can obtain only probabilistically. Each policy has a name, a type, and a configuration block with the same name as the type.
Excluding health checks
Add a policy name: drop-healthchecks. Use type: string_attribute, string_attribute.key: http.route, values of /healthz and /readyz, and invert_match: true.
Every policy is a "condition to keep." So to leave out a specific route, you need an option that flips the match. Use the policy type that matches against a list of attribute values.
An and composite policy and the default probability
Add a policy name: vip-slow. It is type: and, and and.and_sub_policy has two sub-policies — one with type: string_attribute matching tenant.tier equal to enterprise, and one with type: latency and threshold_ms: 300. Finally, add name: baseline (type: probabilistic, probabilistic.sampling_percentage: 2) to make a total of 5 policies.
To keep a trace only when two conditions are met at once, you need a composite policy. The sub-policies are also complete policies, each with a name and a type. At the end, put a probabilistic policy for the rest.
The front-end routing layer
In /root/otca-sampling/loadbalancing.yaml, write the front-end routing layer. In exporters.loadbalancing, set routing_key: traceID, a resolver.dns.hostname that is a service address containing otel-tailsampler, and resolver.dns.port: 4317. The exporter of service.pipelines.traces is just [loadbalancing], and you must not put tail_sampling in the processors.
In front of the decision layer, you need a layer that gathers the spans of the same trace on one instance. An ordinary load balancer will not do; use an exporter that uses the trace ID as its key. This front-end layer does not make decisions.
Buffer memory calculation
In /root/otca-sampling/buffer-memory.txt, write the calculation results on three lines: num_traces_required= (10,000 trace/s × 30s), buffer_mb= (10,000 × 30 × 12 spans × 1.2KB converted to MB and rounded), and buffer_mb_with_headroom= (twice that). Each line is in the 키=값 form (key=value) with no spaces.
It is traces per second times wait time times spans per trace times span size. To convert KB to MB, divide by 1024 and round the decimal. The required value of num_traces comes from multiplying only the first two terms.
OpenTelemetryCollector CR
Create the namespace otca-sampling and write the CR in /root/otca-sampling/collector-cr.yaml. Use apiVersion: opentelemetry.io/v1beta1, kind: OpenTelemetryCollector, metadata.name: otel-tailsampler, metadata.namespace: otca-sampling, spec.mode: deployment, and spec.replicas: 3. Inside spec.config, put the tail_sampling from steps 1–4 (the five policies as they are), and write spec.config.service.pipelines.traces.processors so that tail_sampling comes before batch and batch is last.
In a deployment mode that runs on every node, the spans of one trace are scattered across several nodes and a decision is impossible. The configuration inside the CR has the same structure as a Collector configuration file, and the processor order rule applies as is. Since this is an environment without the CRD, only write the file and actually create only the namespace.