Lakehouse Table Format — Understanding Apache Iceberg Through Its Metadata
Iceberg never deletes anything on its own — compaction, expiration, orphan cleanup
In one line
Because Iceberg only adds files and metadata with every commit, small files are merged by compaction (rewrite_data_files), old snapshots and the files only they used are deleted by expiration (expire_snapshots), and files that no snapshot ever pointed to are cleared by orphan cleanup (remove_orphan_files). All three are irreversible, so order and margin are the answer.
Why cleanup is needed
Picture a table that a streaming load commits to every 5 minutes. In a day, 288 commits, hundreds of small files per partition, 288 snapshots, and 288 metadata.json files pile up. As the Maintenance docs say, small files increase what has to be written in manifests and raise the cost of opening files for every query. And none of this disappears on its own.
It feels odd at first that compaction does not reduce space. Compaction is a new commit that rewrites small files into large ones, and the old small files remain because old snapshots still point to them. That is the very property that made time travel possible. Space comes back after you expire the old snapshots.
How it works — three different jobs
| Job | What it does | Reversible? |
|---|---|---|
| rewrite_data_files | Reads small files, rewrites them as large files, and commits as a replace snapshot | It is a commit, so it can be rolled back |
| expire_snapshots | Removes snapshots older than a threshold from the metadata and deletes the files only they used | No |
| remove_orphan_files | Finds files in the table location that no metadata points to and deletes them | No |
Compaction. The default strategy of rewrite_data_files is binpack, and there is also sort (rewriting while sorting). The target size target-file-size-bytes follows the table property write.target-file-size-bytes (default 512 MB), and files smaller than 75% of the target become candidates for rewriting. min-input-files (default 5) makes a group of that many files get rewritten regardless of other conditions. The result is the replace snapshot the spec talks about — the data of the table is unchanged and only the files change.
Expiration. expire_snapshots deletes snapshots older than older_than (default 5 days ago) but keeps retain_last (default 1) of them. The docs make two things clear. Expiration never deletes files that a snapshot still alive uses, and it does not delete snapshots that a branch or tag points to. So if you put a tag on a point in time you may have to return to, that snapshot and its files survive expiration. Conversely, after expiration you cannot time travel to that snapshot.
Orphan cleanup. When a job fails or loses a commit race, data files remain without entering any snapshot. remove_orphan_files lists the files in the table location and deletes those that the metadata does not point to. There is a trap here. Files being written right now also look like orphans until they are committed. The Maintenance docs warn that if you delete orphans at an interval shorter than the time a write takes to finish, you can delete files in progress and break the table, and they set the default interval to 3 days. The Spark procedure in the lab image goes one step further and rejects an interval shorter than 24 hours altogether. Checking first what would be deleted with dry_run => true should become a habit.
Metadata files pile up too. If you turn on the table property write.metadata.delete-after-commit.enabled (default false), it keeps only write.metadata.previous-versions-max (default 100) and deletes the oldest metadata file on every commit.
What it looks like in the field
The order was reversed. The day after running expiration, you find out there was a reason to go back to the month-end close snapshot. There is no way to undo it. Put "tag before expiration" into the procedure for points in time like closes and deployments.
Orphan cleanup broke a perfectly fine load. Someone said "let's keep only up to yesterday" and shortened older_than to an hour ago, and the files of a load job that happened to be taking long were deleted. If that job succeeds in committing, the table points to files that do not exist. Set the interval with more margin than the longest-running write.
The cleanup jobs shake the table. Compaction is also a commit, so it races with loads. Run it when loads are light, and if it fails, just do it again in the next cycle.
What really matters in practice
- Compaction makes now fast, and expiration gives space back. Compaction alone does not reduce space.
- A tag before expiration. Snapshots that tags and branches point to are not deleted by expiration.
- Do not cut the time margin of orphan cleanup. The default is 3 days, and dry_run first.
- Run the three periodically in a set order. Tag → compaction → expiration → orphan cleanup.
What you will do in the next lab
You load 30 small batches with 30 commits to create a pile of small files and snapshots, and record the numbers before cleanup. After putting a tag on the tenth commit, you compact with rewrite_data_files, and with expire_snapshots you keep only the current snapshot, confirming that the snapshot and files protected by the tag survive. With one old orphan and one just-created file in place, you run remove_orphan_files starting with dry_run and see that only the old one is deleted, and then explain why the number of files on disk after cleanup is larger than the number of files in the current snapshot.