TT Lab
Get started
Learn Learning paths Courses

OTCA — OpenTelemetry Certified Associate

Completing a Full Collector Configuration

Continue in TT Lab

Goal

Write, from start to finish, a Collector configuration file that could go to production. When you are done, you can look at someone else's Collector configuration and predict, from the processor order alone, what kind of incident will happen.

Why it matters

A Collector configuration that is off by even one line of YAML keeps the process from starting, and the problem is that in most cases you learn this only after seeing CrashLoopBackOff in the cluster. Worse is a configuration that starts but is silently wrong. If you put memory_limiter late, there is no problem in normal times and it dies with OOM only under load; if you put batch before tail_sampling, traces are cut at random while the metrics report nothing wrong. If you define a component but do not add it to service, it is in the configuration file but does not run. These three account for most Collector incidents, and all three can be caught just by reading the configuration file.

Steps

  1. Create /root/otca-collector/config.yaml and, under receivers.otlp.protocols, write grpc.endpoint: 0.0.0.0:4317 and http.endpoint: 0.0.0.0:4318.
  2. In processors.memory_limiter, write check_interval: 1s, limit_mib: 1638, and spike_limit_mib: 328. (Based on a container memory limit of 2Gi)
  3. In processors.k8sattributes, write auth_type: serviceAccount, and put four entries in extract.metadata: k8s.namespace.name, k8s.deployment.name, k8s.pod.name, and k8s.node.name.
  4. Create attributes/redact under processors and put three entries in actions: delete for http.request.header.authorization, delete for user.email, and hash for db.query.text.
  5. In processors.batch, write timeout: 5s, send_batch_size: 8192, and send_batch_max_size: 16384.
  6. Define three entries under exporters. otlp/tempo gets an address with port 4317 in endpoint, otlp/gateway also gets an address with port 4317, and prometheusremotewrite gets an http URL ending in /api/v1/push with sending_queue (enabled: true, num_consumers: 10, queue_size: 5000) and retry_on_failure (enabled: true, initial_interval: 5s, max_elapsed_time: 300s) attached.
  7. Define health_check.endpoint: 0.0.0.0:13133 and zpages.endpoint: 0.0.0.0:55679 in extensions, and add both names to service.extensions.
  8. Write three pipelines in service.pipelines. traces uses processors [memory_limiter, k8sattributes, attributes/redact, batch] with exporters [otlp/tempo], metrics uses [memory_limiter, k8sattributes, batch] with [prometheusremotewrite], and logs uses [memory_limiter, k8sattributes, attributes/redact, batch] with [otlp/gateway]. All three pipelines use [otlp] as the receiver.

Notes

Two OTLP receiver ports

Create /root/otca-collector/config.yaml and, under receivers.otlp.protocols, write grpc.endpoint: 0.0.0.0:4317 and http.endpoint: 0.0.0.0:4318.

The otlp receiver has grpc and http separately under protocols. If you open only one, an SDK that does not send to that one fails to connect at all, and no trace is left in the Collector metrics. Inside a container you must bind to all interfaces, not the loopback.

Sizing memory_limiter

In processors.memory_limiter, write check_interval: 1s, limit_mib: 1638, and spike_limit_mib: 328. (Based on a container memory limit of 2Gi)

Assume the container memory limit is 2Gi. The convention is to set the hard limit to 80% of that and the spike allowance to 20% of the hard limit. If you set it higher than the container limit, the kernel kills the process before the processor can step in.

Attaching Kubernetes metadata

In processors.k8sattributes, write auth_type: serviceAccount, and put four entries in extract.metadata: k8s.namespace.name, k8s.deployment.name, k8s.pod.name, and k8s.node.name.

This processor uses the Pod IP as a clue to look up metadata from the API server and attaches it as resource attributes. That is why you must specify the authentication method, and you list only the items you need in extract to reduce load.

Redacting and hashing sensitive attributes

Create attributes/redact under processors and put three entries in actions: delete for http.request.header.authorization, delete for user.email, and hash for db.query.text.

If you add a slash to a processor name, you can create a second instance of the same type. If you delete a value completely, you can no longer group identical values later, so for values that need grouping, such as query text, hash them instead of deleting them.

batch trigger and upper bound

In processors.batch, write timeout: 5s, send_batch_size: 8192, and send_batch_max_size: 16384.

send_batch_size is a trigger meaning "send immediately once this much has accumulated," not an upper bound on batch size. The upper bound must be set with a separate key, and if you do not set it, batches can grow much larger than expected.

Three exporters with queues and retries

Define three entries under exporters. otlp/tempo gets an address with port 4317 in endpoint, otlp/gateway also gets an address with port 4317, and prometheusremotewrite gets an http URL ending in /api/v1/push with sending_queue (enabled: true, num_consumers: 10, queue_size: 5000) and retry_on_failure (enabled: true, initial_interval: 5s, max_elapsed_time: 300s) attached.

sending_queue holds data when the backend slows down briefly, and retry_on_failure retries failures with exponential backoff. If both are off, data is simply lost every time the backend wobbles.

health_check and zpages

Define health_check.endpoint: 0.0.0.0:13133 and zpages.endpoint: 0.0.0.0:55679 in extensions, and add both names to service.extensions.

An extension does not run just because it is defined. You must also add its name under service to activate it. health_check is the endpoint that liveness/readiness probes hit, and zpages shows recent traces and the pipeline state.

Three pipelines and the processor order

Write three pipelines in service.pipelines. traces uses processors [memory_limiter, k8sattributes, attributes/redact, batch] with exporters [otlp/tempo], metrics uses [memory_limiter, k8sattributes, batch] with [prometheusremotewrite], and logs uses [memory_limiter, k8sattributes, attributes/redact, batch] with [otlp/gateway]. All three pipelines use [otlp] as the receiver.

The order you list them in is the processing order. Memory protection must come first so that backpressure goes back upstream, and batch must come last so there is no waste of breaking apart and regrouping. A processor that attaches metadata can create new attributes, so redaction comes after it.