TT Lab
Get started
Learn Learning paths Courses

Finding the Cause in Logs

That Day's Logs Were Already Gone

Continue in TT Lab

One-line summary

Retention is calculation, not intuition. How many days actually remain must be measured by content, not file names, and how many days you should keep must be decided by measuring the compression ratio and comparing it with the disk budget.

Why this is needed

The wall you run into most often on the first day of an investigation is not a hard question but an empty directory. You said "please send me the log from that time three weeks ago," and only five days are left. What you need then is not lament but three numbers — how many days were short, how much more must be kept, and whether that value fits the disk budget.

And all three numbers can be measured from the files that exist now. If you measure one day's original size and the compressed size, you get the compression ratio, and with the compression ratio and one day's volume, you get the bytes for N days. If you know the budget, you get the upper limit of N. There is not one guess in this.

How it works

Rotation has two naming rules. The default of logrotate(8) is the numbering scheme, and there, a larger number is older (app.log.2 is further in the past than app.log.1). If you turn on dateext, a date is attached instead, the default format is -%Y%m%d, and there, a smaller name is older. It is not rare for the two rules to be mixed on one server — the case where only the old files moved to olddir have date names. So "lining files up in time order" becomes the real first step of the investigation.

rotate is a count excluding the current file. The manual defines rotate count as "the number of times files are rotated before being removed," and writes that if count is 0, old files are removed directly rather than rotated. So to keep 30 days, it is daily + rotate 29. An off-by-one mistake is common here.

delaycompress postpones compression by one cycle. The manual states the purpose of this option clearly — because a program that cannot be told to close its log file may keep writing to the previous file for a while. So the most recent two days stay uncompressed, and in capacity calculations those two days must be counted at the original size.

copytruncate loses lines. The same manual states it explicitly — it copies and then truncates the original to 0, and there is a very short gap between the two, so logs written in that gap can be lost. It is a last resort for programs that cannot be restarted or signaled to reopen their files, not something to use as a default.

If you use systemd, the budget is somewhere else. For SystemMaxUse= of journald.conf(5), the default is 10 percent of the filesystem, but capped at 4G. SystemMaxFiles= defaults to 100, and MaxRetentionSec= defaults to 0 — that is, time-based deletion is off. Since data is pushed out by capacity alone, if traffic grows, the retention period quietly gets shorter.

What you see in the field

There are things that do not remain even if you extend retention. RateLimitIntervalSec= and RateLimitBurst= of journald discard the rest of an interval if one service writes more than a set number within a set interval. The default is 10000 per 30 seconds, applied per service, and a message reporting the number discarded is left. The manual adds one more thing here — the effective limit is multiplied according to the free disk space remaining for the journal (a multiplier calculated with a base-2 logarithm). That is, the fuller the disk, the earlier it discards. It is a structure in which the most is discarded at the very moment logs flood during an outage.

So "there is no log" for the incident interval has two meanings. It was deleted by rotation, or it was never recorded in the first place. The remedies are completely different. For the former, you extend retention, and for the latter, you must lower the level, raise the limit, or pull that service's logs out separately.

Compression ratios differ by data. An access log in which similar lines repeat shrinks to under a tenth, but if stack traces or JSON bodies are mixed in, it shrinks far less. So do not use a rough estimate like "usually 10x"; measure with that customer's files. It takes 1 minute to measure, and if you go with a rough guess and the disk fills up, it is not the log but the service that stops.

The thing most often left out of capacity calculations is the uncompressed days. If you use delaycompress, the current file and the previous day's file, two files, are on disk at the original size. If you multiply by the compressed size alone, that much drops out of the budget, and, awkwardly, the better the compression ratio, the larger this error is in relative terms. If one day's original is 200 MB and the compressed copy is 30 MB, the difference of those two days alone is 340 MB — a value that can be a third of a 30-day budget.

And the budget is not used by logs alone. Core dumps, audit logs, and container runtime logs pile up on the same filesystem. That is also why journald's SystemKeepFree= tries to keep 15 percent free by default, and the manual writes that the smaller of the two limits applies. So when setting a rotation policy, you first agree on "how much can be given to logs," and calculate the number of days within that number. If you reverse the order, you fix the number of days you need first and then wait for the disk to fill.

What you will do in the next lab

You build a retention state where rotated copies are mixed with two naming rules, and line them up from oldest. Then you measure by content the interval each file actually holds and show in numbers that the incident interval is out of range. You measure the compression ratio, calculate the bytes for 30 days and the maximum number of days possible within the budget, and write the logrotate settings with that value. Finally, you count the lines that would be discarded under the rate limit setting to show that there is a loss that cannot be prevented by extending retention alone.