TT Lab
Get started
Learn Learning paths Courses

Lakehouse Table Format — Understanding Apache Iceberg Through Its Metadata

Clean up 30 small commits — tag, compact, expire, then orphans, in that order

Continue in TT Lab

Goal

Load 30 small batches with 30 commits to reproduce the pile of small files and snapshots that a streaming load creates, then put a tag on a point in time you may return to, compact with rewrite_data_files, delete old snapshots and their files with expire_snapshots, and clear the files nobody points to with remove_orphan_files. At each step, record how the number of files and the number of snapshots change.

Why it matters

Iceberg deletes nothing on its own. New files and new metadata pile up with every commit, and even when you compact, the old files stay because old snapshots point to them. So a table without cleanup gets slower to read (because of many small files and manifests) and storage keeps growing. Cleanup is three different jobs. Compaction rewrites small files into large ones to make the current snapshot fast. Expiration deletes old snapshots and the files only they pointed to — the past you can reach with time travel shrinks by that much. Orphan cleanup deletes files that no snapshot ever pointed to (traces of failed jobs) — leave enough time margin so that it does not delete even files that are in use right now. All three are irreversible. That is why order matters. If there is a point in time you must return to, you have to tag it before expiration, and you do not cut the time margin of orphan cleanup.

Steps

  1. With /root/ice/mnt/trickle.py (app ice-mnt-trickle), create lake.mnt.orders with PARTITIONED BY (days(order_ts)) and format-version 2, and load /data/ice/batches/batch-001.csv through batch-030.csv with one commit per batch.
  2. With /root/ice/mnt/before.py, write the current number of snapshots, number of data files, and average file size to /root/ice/mnt/out/before.json.
  3. With /root/ice/mnt/tag.py (app ice-mnt-tag), attach the tag batch10 to the snapshot of the tenth commit with RETAIN 30 DAYS.
  4. With /root/ice/mnt/compact.py (app ice-mnt-compact), call rewrite_data_files and write the result to /root/ice/mnt/out/compact.json.
  5. With /root/ice/mnt/expire.py (app ice-mnt-expire), expire snapshots older than now with retain_last => 1.
  6. Put an old orphan (stray-old.parquet, modification time four days ago) and a new file (stray-new.parquet) in /root/ice/warehouse/mnt/orders/data, run remove_orphan_files with /root/ice/mnt/orphans.py (app ice-mnt-orphans) first as a dry_run and write the result to /root/ice/mnt/out/orphans_dry.txt, and then run it for real.
  7. With /root/ice/mnt/after.py, write the number of snapshots, the number of data files, and the number of Parquet files on disk after cleanup to /root/ice/mnt/out/after.json.
  8. In /root/ice/mnt/report.md, write three sections, ## 압축 ## 만료와 태그 ## 고아 파일 (keep these headings as written; they stand for compaction, expiration and tags, and orphan files).

Notes

30 small batches, 30 commits

Create /root/ice/mnt/trickle.py with the app name ice-mnt-trickle, have it create lake.mnt.orders (six columns, PARTITIONED BY (days(order_ts)), 'format-version' = '2'), and append() /data/ice/batches/batch-001.csv through batch-030.csv one at a time in order.

One batch is one commit and one snapshot, and each date partition gets one small file. The grader checks, in the metadata right after the 30th commit, that all 30 snapshots are appends that added the row count of their batch.

The numbers before cleanup

With /root/ice/mnt/before.py (pyiceberg), write the number of snapshots, the number of data files (total-data-files), and the average file size (total-files-size ÷ the number of files, integer division) to /root/ice/mnt/out/before.json as {"snapshots", "data_files", "avg_file_bytes"}.

The summary of the current snapshot holds cumulative values for the whole table (total-*), so you can get the number and size of files without reading all the manifests. If one Parquet file is a few KB, the cost of opening files when reading becomes larger than the cost of reading data.

A name before deleting

Create /root/ice/mnt/tag.py with the app name ice-mnt-tag, have it read lake.mnt.orders.snapshots in committed_at order, and put CREATE TAG batch10 AS OF VERSION <ID> RETAIN 30 DAYS on the tenth snapshot (the placeholder stands for the snapshot ID).

Expiration deletes snapshots older than older_than, but it does not delete snapshots that a tag or branch points to while the retention period of that label remains. So the tag must be attached before expiration. The grader checks that the tag points to the tenth snapshot in the metadata right after the 30th commit.

Compaction — rewrite the small files

Create /root/ice/mnt/compact.py with the app name ice-mnt-compact, call CALL lake.system.rewrite_data_files(table => 'lake.mnt.orders', options => map('min-input-files', '2')), and write the rewritten_data_files_count and added_data_files_count of the result row to /root/ice/mnt/out/compact.json as {"rewritten", "added"}.

Compaction reads the small files of the same partition, rewrites them as larger files, and commits as a single 'replace' snapshot. Not one row changes. The old small files drop out only of the list and stay on disk — because old snapshots (and the tag) still point to them. The grader compares the summary of the replace commit with your two values.

Expiration — old snapshots and their files are deleted

Create /root/ice/mnt/expire.py with the app name ice-mnt-expire, have it read the current time, and call CALL lake.system.expire_snapshots(table => 'lake.mnt.orders', older_than => TIMESTAMP '<지금>', retain_last => 1) (the placeholder stands for the current time).

Expiration removes snapshots from the metadata and deletes from disk the files that none of the remaining snapshots points to. Only the snapshots of the current main and the tag batch10 remain, and the small files of the tenth snapshot that the tag points to are not deleted. The grader looks at the remaining snapshots and the files on disk.

Orphan cleanup — with a time margin

In /root/ice/warehouse/mnt/orders/data, copy one existing data file to make stray-old.parquet (modification time touch -d '4 days ago') and stray-new.parquet (now). Then create /root/ice/mnt/orphans.py with the app name ice-mnt-orphans, write the orphan_file_location of the result of remove_orphan_files(table => 'lake.mnt.orders', dry_run => true) one per line to /root/ice/mnt/out/orphans_dry.txt, and then call it once more without dry_run.

Both files are orphans that no snapshot points to. But the file that was just created might be a file of a commit that someone is writing right now. That is why files newer than older_than are not touched (three days by default). The grader checks that the dry_run list had only the old one and that only the old one was actually deleted.

The numbers after cleanup

With /root/ice/mnt/after.py (pyiceberg), write the number of remaining snapshots, the number of data files in the current snapshot, and the number of Parquet files under /root/ice/warehouse/mnt/orders/data to /root/ice/mnt/out/after.json as {"snapshots", "data_files", "files_on_disk"}.

The number of files on disk is larger than the number of files in the current snapshot. That is because there are the small files that the tag batch10 protects and the new file that orphan cleanup left on purpose. If you can explain that difference, you understand the cleanup jobs.

Cleanup jobs as an operating procedure

In /root/ice/mnt/report.md, write three sections, ## 압축 ## 만료와 태그 ## 고아 파일 (compaction, expiration and tags, and orphan files). In the first section, put the number of data files before compaction (step 2) and the number of data files in the current snapshot after compaction (step 7) as numbers.

If you made these three into jobs that run every day, what order, frequency, and criteria (older_than, retain_last, the orphan time margin) would you set? Also write what would change for a table that is committed to every minute by a stream.