TT Lab
Get started
Learn Learning paths Courses

Finding the Cause in Logs

The Collector Received Fewer Lines Than Were Sent

Continue in TT Lab

Goal

You count loss by comparing the lines the collector received with the sender's counters, fold duplicates, work out the arrival delay distribution, and reveal in numbers the undercount of cutoff-time aggregation and the illusion created by restarts.

Why it matters

When a log is empty, "it did not arrive" and "nothing happened" have opposite conclusions, yet you cannot tell them apart from the file alone. The only thing that separates them is the monotonically increasing number the sender attached. But even that number starts again from 1 when the process restarts, so unless you bind it with a boot identifier, the before and after of a restart look like duplicates and the loss after the restart is hidden. That the order received is not the order that happened is equally important — if you aggregate at the cutoff time, lines that have not yet arrived are missing, and a few days later the same query gives a different number.

Steps

  1. Create and run /root/gap/gen_gap.py to create collector.ndjson and sender_state.json under /root/gap/raw/.
  2. In /root/gap/tally.json, write the result of counting by comparing received lines with sent lines.
  3. In /root/gap/missing.json, write the missing numbers grouped into ranges.
  4. In /root/gap/dedup.ndjson, write the result of folding duplicates.
  5. In /root/gap/delay.json, write the arrival delay distribution.
  6. In /root/gap/cutoff.json, write the undercount of cutoff-time aggregation.
  7. In /root/gap/reboot.json, write the effect of the restart on the aggregation.
  8. Leave a report in four sections in /root/gap/gap_report.md.

Notes

Get what was received and what was sent in hand together

Create and run /root/gap/gen_gap.py to create /root/gap/raw/collector.ndjson (2311 lines, including 61 duplicates) and /root/gap/raw/sender_state.json (4 streams, 2400 sent lines in total).

The collector file piles up in arrival order (observed_ts). Without the sender's counters you can never know what is missing, so you build the two together. One host restarted midway, so the number of (host, boot_id) combinations must be one more than the number of hosts.

Compare received lines with sent lines

In /root/gap/tally.json, write lines_in_file, unique_records, duplicate_lines, sent_total, missing_total, and loss_rate.

The number of lines in the file is not the number of events. You must first fold the same line arriving twice to get 'what was received,' and subtract that from the sender's counters to get 'what did not arrive.' Judge whether it is the same line by (host, boot_id, seq).

Group the missing numbers into ranges

In /root/gap/missing.json, group the missing numbers of each stream into consecutive ranges and write them as host, boot_id, first, last, and count. Sorted by (host, boot_id, first) ascending.

Scattered single-record losses and 50 records dropped as one lump have different causes. So do not just count; group them into ranges. For each stream, sweep from the first_seq to the last_seq the sender reported, and extend the range as long as unreceived numbers continue.

Fold the duplicates created by retransmission

In /root/gap/dedup.ndjson, write the records with duplicates folded as host, boot_id, seq, event_ts, observed_ts, and msg, sorted by (host, boot_id, seq) ascending. If the same line came several times, keep the one that arrived first.

In an at-least-once collector, the same line arriving twice is not a fault but a design. What you need is a definition of 'what to treat as the same line.' If you fold by a fingerprint of the content, even two truly identical events become one, so when a number is available, use the number.

Work out the distribution of arrival delay

In /root/gap/delay.json, write count, p50_sec, p95_sec, max_sec, over_60s, and worst (host, boot_id, seq, delay_sec). The delay is observed_ts minus event_ts truncated to an integer in seconds.

The order received is not the order that happened. This is why you keep the two times separately, and the difference between them is the health of the collection path. The average is of little use — if most are a few seconds and some are minutes, the average explains neither.

How much would you have missed if you aggregated at the cutoff time

In /root/gap/cutoff.json, write cutoff_ts, event_before_cutoff, observed_before_cutoff, late_arrivals, and undercount_rate. The cutoff time is 2026-05-20T03:20:00Z and the comparison is strictly less than.

If you run the same query right after midnight and three days later, the numbers differ. It is not a bug but late arrival. Records whose event time is before the cutoff but whose arrival is after it are exactly the ones you could not count then, and that ratio tells you how much cutoff allowance you need.

A restart swaps loss and duplicates

In /root/gap/reboot.json, write hosts_with_multiple_boots, boots_by_host, false_duplicates, missing_with_boot, and missing_without_boot.

When a process comes back up, the numbers start again from 1. If you count without the boot identifier, the same number before and after the restart looks like a duplicate, and a number missing after the restart is hidden by the pre-restart record and looks like not a loss. Put the two numbers side by side to show that difference.

Write down what to request

In /root/gap/gap_report.md, write four sections: ## 무엇이 얼마나 빠졌나 ## 중복과 늦은 도착 ## 왜 숫자가 달라지나 ## 무엇을 요청할 것인가 (in order, these mean: what is missing and how much, duplicates and late arrivals, why the numbers differ, and what to request). It must include, as numbers, the loss count, the number of duplicate lines, the delay p95, the number not received before the cutoff, and the loss count when boot_id is ignored.

The last section of this report is in practice the most valuable. That is because what you ask them to attach from the next bundle on decides the difficulty of the next investigation. Back up with the earlier numbers why you demand the sender's number, the boot identifier, and both the event time and the observation time.