RDB, AOF, and What You Can Afford to Lose
Summary
A persistence setting is the answer to the question "how many seconds' worth can we afford to lose?". If the answer is "not a single write", you should first reconsider whether Redis is the right tool.
Why this was needed
Redis is an in-memory store. If the process dies, memory is gone. So it provides two persistence methods.
RDB is a periodic snapshot. In a configuration example, save 3600 1, save 300 100, save 60 10000 — it takes a snapshot if at least 1 change occurs within an hour, 100 within 5 minutes, or 10000 within 1 minute. The advantage is that it is a single file, so backup and recovery are simple and restarts are fast. The disadvantage is that you lose all writes since the last snapshot.
AOF logs every write command. appendfsync everysec is the practical default, and in the worst case you lose 1 second's worth. With always, there is almost no loss, but it is an fsync on every write, so throughput drops sharply. no leaves it to the OS, so the loss window grows.
How it works
The practical standard is the hybrid mode that uses both. If you turn on AOF and set aof-use-rdb-preamble yes, then on an AOF rewrite the first part is saved in RDB format and incremental commands are appended after it. Recovery is fast and the loss window stays at 1 second.
Redis 7's Multi-Part AOF tidies this up a step further, separating a base file and incr files in a dedicated directory. The problem of disk space doubling during a rewrite has been reduced.
There is something to point out honestly here. Even with persistence on, Redis is not a database. Replication is asynchronous, so at the moment the master dies there may be writes that did not reach the replica. The WAIT command can improve this partially, but it is not fully synchronous replication. That is why it is not appropriate as the primary store for data that "must never be lost".
What you meet in the field
The cost of fsync is often underestimated. Reports that throughput dropped to less than half after switching to appendfsync always are common. In environments with slow disks, AOF writes can back up and block the main thread.
An RDB snapshot is not free either. It creates a child process with fork to take the snapshot, and because of copy-on-write, memory can grow sharply for a moment when there are many writes. A 4GB instance can end up using 6GB during a snapshot. If you set maxmemory to exactly match physical memory, an OOM occurs at this point. Usually you set maxmemory to about half of physical memory.
The two methods side by side
| RDB (snapshot) | AOF (command log) | |
|---|---|---|
| What it stores | The entire data at that point in time | The commands that changed data, in order |
| File size | Small (compressed binary) | Large (reduced by rewriting) |
| Restart speed | Fast | Slow (re-executes the commands) |
| Worst-case loss | Everything since the last snapshot | Depends on appendfsync |
| Load | Memory spikes at the moment of fork | Writes continuously but gently |
appendfsync is the key knob of AOF.
always 쓸 때마다 fsync — 거의 안 잃지만 아주 느리다
everysec 1초마다 fsync — 최악 1초 유실. 사실상 표준
no OS 에 맡김 — 빠르지만 몇 초를 잃을 수 있다
Turning both on together is the default form. AOF protects you up to the last 1 second, and RDB gives you fast recovery and a backup file. When Redis restarts, if an AOF exists, it uses that first.
Memory spikes caused by fork
An RDB snapshot and an AOF rewrite are done by forking a child process. Thanks to Linux's copy-on-write, memory is shared at first, but while saving, a copy is made for every page the parent changes.
If a snapshot overlaps a moment with many writes, memory rises up to twice as much. This is where incidents happen in which the process hits the container limit and dies with an OOM.
maxmemory 4gb ← 데이터 상한
컨테이너 한도 6~8gb ← fork 여유를 남긴다
This is also why vm.overcommit_memory = 1 is recommended. If this value is 0, the kernel refuses the fork, saying "it looks like memory will run short", and saving itself fails. Redis leaves a warning in the log, but it is easy to overlook.
Decide first what you can afford to lose
- Pure cache — no persistence is needed. If you turn both off, the fork burden also disappears. You can simply refill it.
- Session store — if lost, everyone is logged out. AOF everysec is about right.
- Queues and job state — if lost, jobs disappear. AOF plus a replica.
- Source data — it is better not to use Redis as the source. If you do use it anyway, AOF always with replication, plus regular backups.
The most common mistake is turning persistence on for a cache. You pay for unnecessary fork cost and disk I/O, and then at recovery time an old cache comes back to life and causes problems.
What to look at in the next check
Changing persistence settings could affect the other labs in the lab Pod, so we substitute a quiz. You check whether you can distinguish the loss scope and recovery cost of RDB and AOF by situation, and finish the course.