Three Minutes of the Incident, a Whole Day Read
Goal
From six hours of logs that have been through rotation, you extract just the 3 minutes of the incident interval. You do not open files that do not overlap, in the remaining files you use binary search to mark byte ranges and read only what lies between, and for the compressed copy you sweep from the front and stop early.
Why it matters
Reading a whole day's worth in response to "just show me 3 minutes" is the most common accident in log work. But if the times are fixed in RFC 3339 notation and the file is in time order, string comparison is time comparison, so you can mark the start position of the interval by binary search without reading the file. What is hard is not the search but the boundaries. You must take the interval as half-open so that consecutive intervals do not overlap, and when several lines have the same time, you must find the first line of that time so that the few lines before it do not quietly go missing. And since saying it is fast proves nothing, you must cross-check against the slow method of sweeping through everything once.
Steps
- Create and run
/root/slice/gen_slice.pyto create app.log, app.log.1, and app.log.2.gz under/root/slice/var/log/. /root/slice/index.json— line up the rotated copies from oldest and write the interval each file holds./root/slice/plan.json— pick only the files that overlap the incident interval and skip the rest, with reasons./root/slice/offsets.json— use binary search to mark the start byte and end byte of the interval./root/slice/window_a.log— read only the marked byte ranges and extract the incident interval./root/slice/proof.json— prove that the result is the same by cross-checking against the slow method./root/slice/window_b.logand/root/slice/gz_scan.json— extract the second interval inside the compressed copy, stopping early./root/slice/slice_report.md— leave the method and the grounds as a report.
Notes
- The incident interval is from
2026-04-12T03:58:30.000Zto2026-04-12T04:01:30.000Z, and the second interval is from2026-04-12T00:20:00.000Zto2026-04-12T00:23:00.000Z. Both are half-open intervals that include the start and exclude the end. - The first 24 characters of a line are the time. Since the notation is unified, a string comparison such as
line[:24] >= keyis directly a time comparison. - The spot you
seekto in binary search is almost always in the middle of a line. Read and discard one line to stand on a line boundary, and then judge. - Common mistakes: finding a line that is not the first among lines with the same time, including the line of the end time, leaving out a file beyond the rotation boundary, and applying binary search to the compressed copy.
- Files in the field are several GB, but this lab reproduces the same structure reduced to 20 MB. All outputs are collected under
/root/slice/and disappear when the session ends.
Reproduce a log bundle that has been through rotation
Create and run /root/slice/gen_slice.py to create app.log (57600 lines), app.log.1 (57600 lines), and app.log.2.gz (57600 lines when decompressed) under /root/slice/var/log/.
The three files were written by the same program in the same format and were only split by rotation. Put two hours each at eight lines per second, and make three lines overlap in the same millisecond every 30 seconds. If you write the compressed copy with Python's gzip.GzipFile and give mtime=0, you get the same file even when you recreate it.
Line up the rotated copies in time order and measure the interval they hold
In /root/slice/index.json, write files (an array from oldest, with each item file · compressed · bytes · first_ts · last_ts) and rule (one sentence of the rule by which you lined them up, at least 20 characters). For times, use the first 24 characters of the line as is.
For numbered rotated copies, a larger number is further in the past, and the one with no extension is the file currently being written. If you use logrotate's dateext, the name is a date and the direction is reversed, so also point that out when you write the rule. The last line of an uncompressed file can be pulled out by reading only a few KB from the end of the file — do not read the whole thing. The material is the three files under /root/slice/var/log/.
Do not open files that do not overlap at all
In /root/slice/plan.json, write window (start is 2026-04-12T03:58:30.000Z and end is 2026-04-12T04:01:30.000Z), open (an array of the files that overlap the interval, from oldest), and skip (an array of the remaining files as file · reason).
With just the first_ts and last_ts measured in step 2, you can tell whether a file overlaps without opening it. Since the interval is half-open, the overlap test is half-open too — it overlaps if the file's start is before the interval's end and the file's end is not before the interval's start. The material is /root/slice/index.json. Write reason with at least 10 characters.
Mark the start and end of the interval in bytes with binary search
In /root/slice/offsets.json, write offsets (an array, from oldest, with file · start_offset · end_offset for each uncompressed file picked in step 3). start_offset is the start byte of the first line whose time is at or after the interval start, and end_offset is the start byte of the first line at or after the interval end (the file size if there is none).
If you halve the file size and seek to that byte, you are almost always in the middle of a line. Read and discard one line to stand on a line boundary, then judge by the first 24 characters of that line. There are three lines with the same time, so you must find not 'any line that meets the condition' but 'the line that first meets the condition' — when the condition is met, pull the right end of the range to the start of that line. The material is /root/slice/plan.json.
Read only the marked byte ranges and extract the incident interval
In /root/slice/window_a.log, write the lines of the incident interval [2026-04-12T03:58:30.000Z, 2026-04-12T04:01:30.000Z) in time order, exactly as in the original. Join the files from oldest.
Just read out what lies between the two offsets marked in step 4. end_offset is the start of the first line outside the interval, so everything before it is exactly the half-open interval. Do not reparse or tidy the lines — you must carry over the original bytes as they are for the hash cross-check in the next step to hold. The material is /root/slice/offsets.json.
Prove the result by cross-checking against the slow method
In /root/slice/proof.json, write window, fast_lines, slow_lines, sha256_fast, sha256_slow, match (boolean), bytes_read_fast, and bytes_read_slow. bytes_read_fast is the sum of the bytes between the two offsets in step 4, and bytes_read_slow is the bytes when reading all three files through, including decompressing the compressed copy.
The slow method is to sweep through all three files (decompressing the compressed copy) and collect the lines of the interval. Do not measure time — it wobbles from machine to machine. Instead, compare the line count and the sha256 hash. Compute the hash over the bytes of the extracted lines joined together, and match is true only if the values from the two methods are equal. The materials are /root/slice/window_a.log and /root/slice/offsets.json.
For the compressed copy, stopping early is the answer
The second question is [2026-04-12T00:20:00.000Z, 2026-04-12T00:23:00.000Z). In /root/slice/window_b.log, write the lines of that interval exactly as in the original, and in /root/slice/gz_scan.json, write window, file, total_lines, lines_read, and lines_kept. lines_read is the number of lines actually decompressed until you stopped.
This interval is inside the compressed copy. A gzip stream is decoded depending on the content before it, so you cannot jump to an arbitrary point, and if you apply binary search, you end up decompressing the same spot over and over. Sweep once from the front, and when you meet a line past the end of the interval, get out on the spot. total_lines is a value for writing down how much you saved, so you must count through to the end once. The material is /root/slice/var/log/app.log.2.gz.
Leave the method and the grounds in a document
In /root/slice/slice_report.md, write four sections: ## 무엇을 물었나 ## 왜 통째로 읽지 않아도 되었나 ## 압축본은 무엇이 달랐나 ## 결과를 어떻게 증명했나 (in order, these mean: what was asked, why you did not need to read it all, what was different for the compressed copy, and how the result was proved). It must include, as numbers, the number of lines in the incident interval, the bytes the fast method read, and the number of lines decompressed from the compressed copy.
The person who will read this document is the person who gets asked the same question next time. What that person needs to know is not the commands but the premises — why binary search held, which files you did not open and why, what was different about the compressed copy, and what you cross-checked the result against. Do not make up numbers; take them from /root/slice/proof.json and /root/slice/gz_scan.json.