The Logs from the Day of the Incident Were Already Deleted
Goal
From a bundle of rotated logs, you measure the interval actually retained, show in numbers that the incident interval is outside that range, measure the compression ratio to calculate the retention days and capacity, and write them as a rotation policy. Finally, you count the lines that are never recorded in the first place because of the rate limit.
Why it matters
The wall you run into most often on the first day of an investigation is not a hard question but an empty directory. What you need then is not lament but three numbers — how many days were short, how much more must be kept, and whether that value fits the disk budget. All three numbers can be measured from the files that exist now. And "there is no log" has two meanings. It was deleted by rotation, or it was never recorded in the first place because of the rate limit. The remedies are completely different, so you must tell the two apart.
Steps
- Create and run
/root/keep/gen_keep.pyto reproduce the retention state under/root/keep/var/log/. - In
/root/keep/order.json, write the list of rotated copies lined up from oldest and its rule. - In
/root/keep/coverage.json, write each file's line count and the times of its first and last lines. - In
/root/keep/window.json, write whether the incident interval is within the retained range and by how much it fell short. - In
/root/keep/budget.json, write the measured compression ratio and the retention days and capacity calculation. - In
/root/keep/logrotate.conf, write the rotation policy. Therotatevalue must be the same as the calculation in step 5. - In
/root/keep/ratelimit.json, count and write the number of lines that would be discarded by the rate limit. - Leave a report in four sections in
/root/keep/keep_report.md.
Notes
- These are the assumptions this lab gives. The incident interval is from
2026-04-05T02:10:00Zto2026-04-05T02:40:00Z. The production server produces 1200 times this sample. The disk budget for/var/logis 4 GiB (4294967296 bytes), the required retention period is 30 days, the journald settings areRateLimitIntervalSec=30andRateLimitBurst=200, and the multiplier by free space is taken as 1. - The calculation rules for step 5.
sample_raw_bytes_per_dayis the size ofpayments.log.1, andsample_gz_bytes_per_dayis the size ofpayments.log.2.gz.compression_ratiois the ratio of the two, rounded to the fourth decimal place. Because ofdelaycompress, the most recent two days are treated as uncompressed.required_bytes_for_30_daysis two days of original + 28 days compressed, andmax_days_in_budgetis the quotient of the budget minus two days of original, divided by one day of compressed, plus 2.rotate_valueis the number you write in logrotate when keeping 30 days. - You can read a compressed copy without decompressing it using
zcatandzgrep, and in Python it isgzip.open(path, "rt"). - Common mistakes: writing
rotateequal to the retention days, sorting numbered names and date names in the same direction, putting in the compression ratio as a rough estimate, and calculating the most recent two days at the compressed size. - Collect all outputs under
/root/keep/. They disappear when the session ends.
Reproduce the customer's retention state
Create and run /root/keep/gen_keep.py to create payments.log (720 lines), payments.log.1 (2000 lines), payments.log.2.gz, payments.log.3.gz, old/payments.log-20260407.gz, old/payments.log-20260408.gz, and burst.ndjson (360 lines) under /root/keep/var/log/.
You build a state where rotation names are mixed in two ways. This lab holds only if numbered and date-named files, and compressed and uncompressed files, are all present together. If you write the compressed copies with Python's gzip.GzipFile and give mtime=0, you get the same file even when you recreate it.
Line up the rotated copies from oldest
In /root/keep/order.json, write order (an array of file paths from oldest, as relative paths with /root/keep/var/log/ removed) and rule (one sentence of the rule by which you lined them up).
Numbered names and date names sort in opposite directions. On one side a larger number is further in the past, and on the other a smaller name is further in the past. The file currently being written is the newest. Line them up by name alone without opening the files, and confirm by content in the next step.
Measure the interval each file actually holds
In /root/keep/coverage.json, write, as an array from oldest, an object for each file containing file, lines, first_ts, and last_ts. For times, use the first field of the log line as is.
File names can lie — if rotation failed or a file moved by hand gets mixed in, the name order and the content order diverge. So you measure again by content. Read a compressed copy with zcat, or open it in Python with gzip.open(path, 'rt').
Show in numbers that the incident interval is out of range
In /root/keep/window.json, write oldest_retained, newest_retained, incident_start, incident_end, incident_covered (true/false), short_by_seconds (integer), and extra_rotations_needed (integer). extra_rotations_needed is the seconds short divided by one day and rounded up.
The incident interval is written in the notes section of the instructions. The seconds short are the incident start time subtracted from the oldest retained time. Since rotation is by day, you find how many more days you should have kept by rounding up.
Measure the compression ratio and calculate the retention days
In /root/keep/budget.json, write sample_raw_bytes_per_day, sample_gz_bytes_per_day, compression_ratio, prod_raw_bytes_per_day, prod_gz_bytes_per_day, required_bytes_for_30_days, fits_in_budget, max_days_in_budget, and rotate_value according to the calculation rules in the instructions.
Do not use a rough estimate for the compression ratio; measure it with this customer's files — access logs and stack traces shrink to different degrees. Because of delaycompress, the most recent two days must be counted at the original size, and that logrotate's rotate is a count excluding the current file is the trap in this calculation.
Write the rotation policy with the calculated value
In /root/keep/logrotate.conf, write a /root/keep/var/log/payments.log block. It must have daily · rotate <5단계의 rotate_value> (where the placeholder is the rotate_value from step 5) · compress · delaycompress · missingok · notifempty · dateext · create 0640 root adm, and do not include copytruncate.
The rotate value is the number you calculated in step 5 — it is not the same as the retention days. The reason to leave out copytruncate is written in the manual. It is because lines written in the short gap between copying and truncating vanish, and they vanish at the moment you need them most in an investigation.
Count the lines that are never recorded in the first place
In /root/keep/ratelimit.json, write interval_sec, burst, windows_over_limit, dropped_total, dropped_by_service (an object keyed by service name), and worst_window (window_start, service, dropped). The lines discarded in one interval are count - burst, and 0 if negative.
journald's rate limit is applied separately per service, and if the limit is exceeded in an interval, the entire rest of that interval is discarded. However much you extend the retention period, this loss remains. worst_window is the one interval with the most discarded lines.
Write down what to change
In /root/keep/keep_report.md, write four sections: ## 무엇이 없었나 ## 지금 보관 정책은 무엇인가 ## 얼마를 남겨야 하는가 ## 무엇을 바꾸기로 했나 (in order, these mean: what was missing, what the current retention policy is, how much must be kept, and what we decided to change). It must include, as numbers, the seconds short, the rotate value, the number of days possible within the budget, and the number of lines that would be discarded by the rate limit.
The person who reads this report is the person who holds the disk budget. A decision comes only if you say not 'please keep more logs' but 'N bytes a day, M bytes for 30 days, it fits within the budget.' Also write that the rate-limit loss needs a different remedy from retention.