create and copytruncate
In one line
create does not lose logs but needs the process to cooperate, and copytruncate needs no cooperation but loses a few logs. Neither is free.
Why this was needed
When a log file grows, the disk fills up. But the process keeps the file open while it writes, so you cannot just delete it. On Linux, rm removes only the directory entry, and the blocks of an open file are not returned until the last handle is closed. So df says the disk is full, while du says the file isn't there.
Rotation solves this problem in two ways.
create — rename and create a new file
app.log → app.log.1 로 rename
새 app.log 생성
프로세스에 SIGHUP 을 보내 다시 열게 한다
Since rename leaves the inode as it is, the process still writes to the old file. That is why you must make it reopen the file. If you don't, the new app.log stays at 0 bytes forever, and the real logs keep flowing into app.log.1. You don't lose logs, but nobody can find them.
copytruncate — copy, then truncate
app.log → app.log.1 로 복사
app.log 를 0바이트로 truncate
Because the file stays the same, the process doesn't need to know anything. In exchange, logs written between the copy and the truncate are lost. For a log that piles up thousands of lines per second, that gap is large.
How it works
/var/log/app/*.log {
daily
rotate 14 # 14개까지 보관
size 100M # 또는 100MB 넘으면
compress
delaycompress # 직전 것은 압축하지 않는다(아직 쓰고 있을 수 있다)
missingok
notifempty
create 0640 app app
postrotate
kill -HUP $(cat /run/app.pid) 2>/dev/null || true
endscript
}
delaycompress is important. With the create method, the process may still be writing to the old file, and compressing it right away breaks that write.
Common misconceptions
Running logrotate in a container. Containers usually write logs to stdout. Then the container runtime (containerd) receives them into a file and the runtime does the rotation. The kubelet's containerLogMaxSize (default 10Mi) and containerLogMaxFiles (default 5) are those settings. There is no reason to run logrotate inside a container.
Rotation breaking the collector. A collector such as Fluent Bit tracks files by inode. When a file is renamed by create, the collector may keep reading the old inode and miss the new file. The Rotate_Wait setting fills that gap.
Don't rotate yourself in containers
Everything so far applies to processes that write to files. In a container, you write to standard output and leave the rotation to the runtime.
// /etc/docker/daemon.json 또는 containerd 설정
{"log-driver": "json-file",
"log-opts": {"max-size": "10m", "max-file": "3"}}
Without this setting, log files grow without limit and fill the node's disk. Then
every Pod on that node dies. It is a common cause of the incident in which Kubernetes evicts Pods from a node under DiskPressure.
The kubelet has the same settings (containerLogMaxSize, containerLogMaxFiles).
Which one applies, the container runtime setting or the kubelet setting, depends on the runtime,
so the reliable approach is to check the actual file sizes on the node.
du -sh /var/log/pods/* | sort -h | tail -5
Three places where logs disappear
Apart from rotation, logs quietly vanish in several places.
Buffering. When a process dies, whatever was in its buffer is lost. Switch to line-by-line output with
PYTHONUNBUFFERED=1 for Python and stdbuf -oL for shell scripts.
Collector queue overflow. Fluent Bit and Vector drop new logs when the memory buffer fills up (the default policy). Switching to a disk buffer reduces loss but uses the node's disk.
# Fluent Bit — 넘칠 때 디스크로
[SERVICE]
storage.path /var/log/flb-storage/
storage.max_chunks_up 128
[INPUT]
storage.type filesystem
Container deletion. When a Pod is deleted, everything under /var/log/pods disappears too. So to
see a dead Pod's logs later, the collector must have taken them beforehand. If the collector
scans every 30 seconds and the Pod is deleted in between, the last logs are gone. Raising terminationGrace PeriodSeconds a little narrows that window.
What to log
However well you configure rotation, a large volume becomes a cost. Following three rules cuts it down a lot.
- Summarize the normal path. One line per request is enough. Leaving debug logs on in production is the number one cause of log volume explosions.
- Aggregate what repeats. If the same error occurs 1,000 times per second, one line saying "this error, 1,000 times" is better than 1,000 lines.
- Structure them. If you write JSON, you can filter by field later, so you can divide retention policies more finely.
What really matters in practice
Set the retention period by regulation and investigation needs, not by disk. Instead of "there's disk left, so 90 days," if "incident investigations usually need 2 weeks and the audit requirement is 1 year," split it into hot storage for 2 weeks plus cold storage for 1 year.
And compression is almost always a win. Text logs shrink 10–20 times with gzip. It costs a little CPU and saves a lot of disk.