TT Lab
Get started
Learn Learning paths Courses

OTCA — OpenTelemetry Certified Associate

Tail Sampling Policy and the Routing Layer

Continue in TT Lab

Goal

Design five tail sampling policies, put a trace ID routing layer in front of them, calculate the buffer memory yourself, and reflect that value in an Operator CR.

Why it matters

Tail sampling is not a feature where "turn it on and cost goes down." To make a decision, every span must arrive, so the network from the agent to the Collector is not reduced at all, and what shrinks is only backend storage and indexing. In exchange, you pay three new things: the memory to hold every span during decision_wait, the routing layer that gathers the same trace on one instance, and the complexity of operating that layer. So the decision to adopt it is made by multiplication, not by feel. Multiply traces per second by the wait time to get num_traces, and multiply that by the span count and size to get the memory. If you skip this calculation, you run with too small a limit, old traces are forced to a decision, and you meet the situation where data disappears while the metrics report nothing wrong.

Steps

  1. Create /root/otca-sampling/tailsampling.yaml and, in processors.tail_sampling, write decision_wait: 30s, num_traces: 300000, and expected_new_traces_per_sec: 10000.
  2. Put two policies in policies: name: keep-errors (type: status_code, with ERROR in status_code.status_codes) and name: keep-slow (type: latency, latency.threshold_ms: 800).
  3. Add a policy name: drop-healthchecks. Use type: string_attribute, string_attribute.key: http.route, values of /healthz and /readyz, and invert_match: true.
  4. Add a policy name: vip-slow. It is type: and, and and.and_sub_policy has two sub-policies — one with type: string_attribute matching tenant.tier equal to enterprise, and one with type: latency and threshold_ms: 300. Finally, add name: baseline (type: probabilistic, probabilistic.sampling_percentage: 2) to make a total of 5 policies.
  5. In /root/otca-sampling/loadbalancing.yaml, write the front-end routing layer. In exporters.loadbalancing, set routing_key: traceID, a resolver.dns.hostname that is a service address containing otel-tailsampler, and resolver.dns.port: 4317. The exporter of service.pipelines.traces is just [loadbalancing], and you must not put tail_sampling in the processors.
  6. In /root/otca-sampling/buffer-memory.txt, write the calculation results on three lines: num_traces_required= (10,000 trace/s × 30s), buffer_mb= (10,000 × 30 × 12 spans × 1.2KB converted to MB and rounded), and buffer_mb_with_headroom= (twice that). Each line is in the 키=값 form (key=value) with no spaces.
  7. Create the namespace otca-sampling and write the CR in /root/otca-sampling/collector-cr.yaml. Use apiVersion: opentelemetry.io/v1beta1, kind: OpenTelemetryCollector, metadata.name: otel-tailsampler, metadata.namespace: otca-sampling, spec.mode: deployment, and spec.replicas: 3. Inside spec.config, put the tail_sampling from steps 1–4 (the five policies as they are), and write spec.config.service.pipelines.traces.processors so that tail_sampling comes before batch and batch is last.

Notes

tail_sampling defaults

Create /root/otca-sampling/tailsampling.yaml and, in processors.tail_sampling, write decision_wait: 30s, num_traces: 300000, and expected_new_traces_per_sec: 10000.

decision_wait is the time margin for all spans of a trace to arrive. If it is too short, you decide on an incomplete trace, and if it is too long, memory spikes. num_traces comes from the product of traces per second and the wait time.

Two outcome-based policies

Put two policies in policies: name: keep-errors (type: status_code, with ERROR in status_code.status_codes) and name: keep-slow (type: latency, latency.threshold_ms: 800).

These two policies are the reason tail sampling exists. They retain 100% by rule what head sampling can obtain only probabilistically. Each policy has a name, a type, and a configuration block with the same name as the type.

Excluding health checks

Add a policy name: drop-healthchecks. Use type: string_attribute, string_attribute.key: http.route, values of /healthz and /readyz, and invert_match: true.

Every policy is a "condition to keep." So to leave out a specific route, you need an option that flips the match. Use the policy type that matches against a list of attribute values.

An and composite policy and the default probability

Add a policy name: vip-slow. It is type: and, and and.and_sub_policy has two sub-policies — one with type: string_attribute matching tenant.tier equal to enterprise, and one with type: latency and threshold_ms: 300. Finally, add name: baseline (type: probabilistic, probabilistic.sampling_percentage: 2) to make a total of 5 policies.

To keep a trace only when two conditions are met at once, you need a composite policy. The sub-policies are also complete policies, each with a name and a type. At the end, put a probabilistic policy for the rest.

The front-end routing layer

In /root/otca-sampling/loadbalancing.yaml, write the front-end routing layer. In exporters.loadbalancing, set routing_key: traceID, a resolver.dns.hostname that is a service address containing otel-tailsampler, and resolver.dns.port: 4317. The exporter of service.pipelines.traces is just [loadbalancing], and you must not put tail_sampling in the processors.

In front of the decision layer, you need a layer that gathers the spans of the same trace on one instance. An ordinary load balancer will not do; use an exporter that uses the trace ID as its key. This front-end layer does not make decisions.

Buffer memory calculation

In /root/otca-sampling/buffer-memory.txt, write the calculation results on three lines: num_traces_required= (10,000 trace/s × 30s), buffer_mb= (10,000 × 30 × 12 spans × 1.2KB converted to MB and rounded), and buffer_mb_with_headroom= (twice that). Each line is in the 키=값 form (key=value) with no spaces.

It is traces per second times wait time times spans per trace times span size. To convert KB to MB, divide by 1024 and round the decimal. The required value of num_traces comes from multiplying only the first two terms.

OpenTelemetryCollector CR

Create the namespace otca-sampling and write the CR in /root/otca-sampling/collector-cr.yaml. Use apiVersion: opentelemetry.io/v1beta1, kind: OpenTelemetryCollector, metadata.name: otel-tailsampler, metadata.namespace: otca-sampling, spec.mode: deployment, and spec.replicas: 3. Inside spec.config, put the tail_sampling from steps 1–4 (the five policies as they are), and write spec.config.service.pipelines.traces.processors so that tail_sampling comes before batch and batch is last.

In a deployment mode that runs on every node, the spans of one trace are scattered across several nodes and a decision is impossible. The configuration inside the CR has the same structure as a Collector configuration file, and the processor order rule applies as is. Since this is an environment without the CRD, only write the file and actually create only the namespace.