The Collector Received Fewer Lines Than Were Sent
Goal
You count loss by comparing the lines the collector received with the sender's counters, fold duplicates, work out the arrival delay distribution, and reveal in numbers the undercount of cutoff-time aggregation and the illusion created by restarts.
Why it matters
When a log is empty, "it did not arrive" and "nothing happened" have opposite conclusions, yet you cannot tell them apart from the file alone. The only thing that separates them is the monotonically increasing number the sender attached. But even that number starts again from 1 when the process restarts, so unless you bind it with a boot identifier, the before and after of a restart look like duplicates and the loss after the restart is hidden. That the order received is not the order that happened is equally important — if you aggregate at the cutoff time, lines that have not yet arrived are missing, and a few days later the same query gives a different number.
Steps
- Create and run
/root/gap/gen_gap.pyto create collector.ndjson and sender_state.json under/root/gap/raw/. - In
/root/gap/tally.json, write the result of counting by comparing received lines with sent lines. - In
/root/gap/missing.json, write the missing numbers grouped into ranges. - In
/root/gap/dedup.ndjson, write the result of folding duplicates. - In
/root/gap/delay.json, write the arrival delay distribution. - In
/root/gap/cutoff.json, write the undercount of cutoff-time aggregation. - In
/root/gap/reboot.json, write the effect of the restart on the aggregation. - Leave a report in four sections in
/root/gap/gap_report.md.
Notes
- The key names and calculation rules of the outputs.
tally.json— lines_in_file · unique_records · duplicate_lines · sent_total · missing_total · loss_rate (loss divided by sent lines, rounded to the fourth decimal place).missing.json— an array, each item host · boot_id · first · last · count (both ends inclusive), sorted by (host, boot_id, first) ascending.dedup.ndjson— host · boot_id · seq · event_ts · observed_ts · msg, sorted by (host, boot_id, seq) ascending. If the same line came several times, keep the one that arrived first (the one with the earliest observed_ts).delay.json— count · p50_sec · p95_sec · max_sec · over_60s · worst (host · boot_id · seq · delay_sec). The delay is observed_ts minus event_ts truncated to an integer in seconds, and a percentile is theceil(p/100 × n)-th value after sorting (counting from 1). over_60s is the number exceeding 60 seconds.cutoff.json— cutoff_ts · event_before_cutoff · observed_before_cutoff · late_arrivals · undercount_rate (late arrivals divided by event_before_cutoff, rounded to the fourth decimal place). The cutoff time is2026-05-20T03:20:00Zand the comparison is strictly less than.reboot.json— hosts_with_multiple_boots (ascending array) · boots_by_host (host as key, value an ascending array) · false_duplicates · missing_with_boot · missing_without_boot. - false_duplicates is the number of unique counts by (host, boot_id, seq) minus the number of unique counts by (host, seq) alone. missing_without_boot is, for each host, the number of numbers not received from 1 up to that host's largest number, ignoring boots.
- Steps 4, 5, 6, and 7 count based on the records with duplicates folded.
- Common mistakes: counting without folding duplicates, counting late-arriving lines as loss, judging the same line by seq alone without boot_id, and rounding the delay.
- Collect all outputs under
/root/gap/. They disappear when the session ends.
Get what was received and what was sent in hand together
Create and run /root/gap/gen_gap.py to create /root/gap/raw/collector.ndjson (2311 lines, including 61 duplicates) and /root/gap/raw/sender_state.json (4 streams, 2400 sent lines in total).
The collector file piles up in arrival order (observed_ts). Without the sender's counters you can never know what is missing, so you build the two together. One host restarted midway, so the number of (host, boot_id) combinations must be one more than the number of hosts.
Compare received lines with sent lines
In /root/gap/tally.json, write lines_in_file, unique_records, duplicate_lines, sent_total, missing_total, and loss_rate.
The number of lines in the file is not the number of events. You must first fold the same line arriving twice to get 'what was received,' and subtract that from the sender's counters to get 'what did not arrive.' Judge whether it is the same line by (host, boot_id, seq).
Group the missing numbers into ranges
In /root/gap/missing.json, group the missing numbers of each stream into consecutive ranges and write them as host, boot_id, first, last, and count. Sorted by (host, boot_id, first) ascending.
Scattered single-record losses and 50 records dropped as one lump have different causes. So do not just count; group them into ranges. For each stream, sweep from the first_seq to the last_seq the sender reported, and extend the range as long as unreceived numbers continue.
Fold the duplicates created by retransmission
In /root/gap/dedup.ndjson, write the records with duplicates folded as host, boot_id, seq, event_ts, observed_ts, and msg, sorted by (host, boot_id, seq) ascending. If the same line came several times, keep the one that arrived first.
In an at-least-once collector, the same line arriving twice is not a fault but a design. What you need is a definition of 'what to treat as the same line.' If you fold by a fingerprint of the content, even two truly identical events become one, so when a number is available, use the number.
Work out the distribution of arrival delay
In /root/gap/delay.json, write count, p50_sec, p95_sec, max_sec, over_60s, and worst (host, boot_id, seq, delay_sec). The delay is observed_ts minus event_ts truncated to an integer in seconds.
The order received is not the order that happened. This is why you keep the two times separately, and the difference between them is the health of the collection path. The average is of little use — if most are a few seconds and some are minutes, the average explains neither.
How much would you have missed if you aggregated at the cutoff time
In /root/gap/cutoff.json, write cutoff_ts, event_before_cutoff, observed_before_cutoff, late_arrivals, and undercount_rate. The cutoff time is 2026-05-20T03:20:00Z and the comparison is strictly less than.
If you run the same query right after midnight and three days later, the numbers differ. It is not a bug but late arrival. Records whose event time is before the cutoff but whose arrival is after it are exactly the ones you could not count then, and that ratio tells you how much cutoff allowance you need.
A restart swaps loss and duplicates
In /root/gap/reboot.json, write hosts_with_multiple_boots, boots_by_host, false_duplicates, missing_with_boot, and missing_without_boot.
When a process comes back up, the numbers start again from 1. If you count without the boot identifier, the same number before and after the restart looks like a duplicate, and a number missing after the restart is hidden by the pre-restart record and looks like not a loss. Put the two numbers side by side to show that difference.
Write down what to request
In /root/gap/gap_report.md, write four sections: ## 무엇이 얼마나 빠졌나 ## 중복과 늦은 도착 ## 왜 숫자가 달라지나 ## 무엇을 요청할 것인가 (in order, these mean: what is missing and how much, duplicates and late arrivals, why the numbers differ, and what to request). It must include, as numbers, the loss count, the number of duplicate lines, the delay p95, the number not received before the cutoff, and the loss count when boot_id is ignored.
The last section of this report is in practice the most valuable. That is because what you ask them to attach from the next bundle on decides the difficulty of the next investigation. Back up with the earlier numbers why you demand the sender's number, the boot identifier, and both the event time and the observation time.