OTCA — OpenTelemetry Certified Associate
The Collector returned 200, but the spans were nowhere
Goal
Measure for real how the response the caller receives and the internal telemetry metrics change depending on the Collector's queue and retry settings when the backend stalls or rejects requests, and design a configuration that gets through a short outage without loss.
Why it matters
An application receiving a 200 does not mean the spans reached storage. A queue decouples the caller from the outage but moves the loss inside the Collector, retries save only transient failures, and past the limit the data disappears, leaving only one log line and one metric. You can diagnose pipeline loss only if you can read otelcol_receiver_accepted_spans, otelcol_exporter_sent_spans, send_failed, and enqueue_failed side by side.
Prepared environment
python3 /opt/fixtures/otca_backpressure_lab.py init places broken.yaml and spans.json (2 spans per request) in /root/otca-export/. python3 /opt/fixtures/otca_backpressure_lab.py run 설정.yaml [--backend down|ok|reject] [--backend-after 초] [--requests N] [--wait 초] (the placeholders are the configuration file and the number of seconds) actually starts lab-k8s's otelcol-contrib 0.116.0, sends requests at 0.2-second intervals, waits, and then reads the otelcol_* metrics from the Collector's internal telemetry and summarizes them. The endpoint of the otlphttp exporter is swapped for an experimental receiver (down is a port nobody listens on, ok returns 200, and reject returns 400). The grader reruns your configuration under the same conditions and compares it with the numbers you wrote.
Steps
- Create the materials with
python3 /opt/fixtures/otca_backpressure_lab.py init, then runotelcol-contrib validate --config /root/otca-export/broken.yaml. In/root/otca-export/01-validate.txt, writemissing_component=(the nonexistent component the error points to), and create/root/otca-export/fixed.yaml, fixed so that the pipeline exports to the definedotlphttp/backend. - Create
/root/otca-export/no-queue.yamlfrom fixed.yaml and putsending_queue.enabled: falseandretry_on_failure.enabled: falseinotlphttp/backend. Transfer the result ofpython3 /opt/fixtures/otca_backpressure_lab.py run /root/otca-export/no-queue.yaml(backend down, 5 requests × 2 spans) into/root/otca-export/02-no-queue.txtascodes=,accepted_spans=,refused_spans=, andsend_failed_spans=. - In
/root/otca-export/queue.yaml, turn onsending_queueandretry_on_failure(other values at their defaults). Run it the same way (backend down) and writecodes=,accepted_spans=,sent_spans=, andsend_failed_spans=into/root/otca-export/03-queue.txt. - In
/root/otca-export/small-queue.yaml, putsending_queue: {enabled: true, queue_size: 2, num_consumers: 1}andretry_on_failure: {enabled: true, initial_interval: 1s, max_interval: 1s}. Run it with the backend down and writecodes=,accepted_spans=,refused_spans=,enqueue_failed_spans=, andqueue_size=into/root/otca-export/04-full.txt. - Run the same small-queue.yaml with
--backend ok --backend-after 2.5 --wait 6(the receiver comes up after 2.5 seconds). Writeaccepted_spans=,sent_spans=, andsend_failed_spans=into/root/otca-export/05-recover.txt. - Run queue.yaml with
--backend reject(the receiver returns HTTP 400). In/root/otca-export/06-permanent.txt, writecodes=,sent_spans=,send_failed_spans=,dropping_logged=(true/false for whether Dropping data appeared in the log), andretried=(true/false for whether the number of requests the receiver got, backend_requests, was more than the 5 requests). - In
/root/otca-export/short-retry.yaml, setretry_on_failuretoinitial_interval: 500ms,max_interval: 500ms, andmax_elapsed_time: 2s(queue at its defaults). Run it with the backend down and--wait 6, and writeaccepted_spans=,sent_spans=,send_failed_spans=, anddropping_logged=into/root/otca-export/07-give-up.txt. - Design
/root/otca-export/resilient.yaml. The grader runs it under the conditions--backend ok --backend-after 4 --requests 8 --wait 8and checks that there are 0 rejections to the caller, all 16 spans are sent within 8 seconds, and there are 0 failures. First check withpython3 /opt/fixtures/otca_backpressure_lab.py run /root/otca-export/resilient.yaml --backend ok --backend-after 4 --requests 8 --wait 8.
Notes
- Accepted and refused are receiver metrics; sent, send_failed, and enqueue_failed are exporter metrics.
- Common mistakes: reading a 200 response as delivery complete, thinking that making the queue bigger eliminates loss (retry limits, memory, restarts), and expecting a 400 to be retried too.
- The queue in this lab is an in-memory queue, so it is emptied when the Collector restarts. A persistent queue (the storage extension) is not covered. The numbers are values measured under this version and short experimental conditions.
- Exporter helper · Internal telemetry
An error that could have been caught before startup
Create the materials with python3 /opt/fixtures/otca_backpressure_lab.py init, then run otelcol-contrib validate --config /root/otca-export/broken.yaml. In /root/otca-export/01-validate.txt, write missing_component= (the nonexistent component the error points to), and create /root/otca-export/fixed.yaml, fixed so that the pipeline exports to the defined otlphttp/backend.
validate checks the reference relationships in the configuration without starting the Collector. Compare the names the pipeline calls with the names defined in exporters.
With neither queue nor retry
Create /root/otca-export/no-queue.yaml from fixed.yaml and put sending_queue.enabled: false and retry_on_failure.enabled: false in otlphttp/backend. Transfer the result of python3 /opt/fixtures/otca_backpressure_lab.py run /root/otca-export/no-queue.yaml (backend down, 5 requests × 2 spans) into /root/otca-export/02-no-queue.txt as codes=, accepted_spans=, refused_spans=, and send_failed_spans=.
Without a queue, the exporter call happens synchronously inside the receiver request. Check the response code to see how far back the failure travels.
A 200 came back, but nothing was sent
In /root/otca-export/queue.yaml, turn on sending_queue and retry_on_failure (other values at their defaults). Run it the same way (backend down) and write codes=, accepted_spans=, sent_spans=, and send_failed_spans= into /root/otca-export/03-queue.txt.
With a queue, the receiver returns success the moment it puts the data in the queue. Accepted and delivered are different metrics.
When the queue fills, refusals come back
In /root/otca-export/small-queue.yaml, put sending_queue: {enabled: true, queue_size: 2, num_consumers: 1} and retry_on_failure: {enabled: true, initial_interval: 1s, max_interval: 1s}. Run it with the backend down and write codes=, accepted_spans=, refused_spans=, enqueue_failed_spans=, and queue_size= into /root/otca-export/04-full.txt.
While one consumer holds the first batch and retries, the queue has only two slots. Requests arrive at 0.2-second intervals.
What survives when the backend comes back
Run the same small-queue.yaml with --backend ok --backend-after 2.5 --wait 6 (the receiver comes up after 2.5 seconds). Write accepted_spans=, sent_spans=, and send_failed_spans= into /root/otca-export/05-recover.txt.
What got into the queue survives through retries, but what could not get into the queue and was refused is not held by the Collector. Resending is the caller's job.
A 400 is not resent
Run queue.yaml with --backend reject (the receiver returns HTTP 400). In /root/otca-export/06-permanent.txt, write codes=, sent_spans=, send_failed_spans=, dropping_logged= (true/false for whether Dropping data appeared in the log), and retried= (true/false for whether the number of requests the receiver got, backend_requests, was more than the 5 requests).
Retry is a mechanism for transient failures (connection refused, 503, and so on). A response saying the format is wrong stays the same no matter how many times you send it.
Retries have an end too
In /root/otca-export/short-retry.yaml, set retry_on_failure to initial_interval: 500ms, max_interval: 500ms, and max_elapsed_time: 2s (queue at its defaults). Run it with the backend down and --wait 6, and write accepted_spans=, sent_spans=, send_failed_spans=, and dropping_logged= into /root/otca-export/07-give-up.txt.
When max_elapsed_time passes, retries stop and that batch is discarded. The caller has already received a 200, so the loss is left only in the Collector log and internal metrics.
A configuration that survives a 4-second outage
Design /root/otca-export/resilient.yaml. The grader runs it under the conditions --backend ok --backend-after 4 --requests 8 --wait 8 and checks that there are 0 rejections to the caller, all 16 spans are sent within 8 seconds, and there are 0 failures. First check with python3 /opt/fixtures/otca_backpressure_lab.py run /root/otca-export/resilient.yaml --backend ok --backend-after 4 --requests 8 --wait 8.
Two things are needed — room to hold all the batches that arrive during the outage (a batch being retried is held by a consumer, so look at queue_size and num_consumers together for the room), and an interval such that the first retry arrives within the wait time after the backend returns. Check the default first-retry interval in the documentation.