TT Lab
Get started
Learn Learning paths Courses

Lakehouse Table Format — Understanding Apache Iceberg Through Its Metadata

A table is a file list, not a directory — from metadata.json down to data files

Continue in TT Lab

In one line

An Iceberg table is not found by scanning a directory. It is defined by a tree of lists — metadata.json → manifest list → manifest → data file — and a commit means writing a new tree and then atomically swapping a single pointer in the catalog.

Why directories were not enough

In a Hive-style table, the "table" is a directory. The files under orders/order_date=2026-03-01/ are that day's data, and the reading engine lists the directory to collect them. This is simple, but it has three problems.

First, there is no atomicity. If a reader lists the directory when a writing job has uploaded five of its ten files, the reader sees a half-finished result. Second, listing is expensive. As the Reliability docs point out, a Hive table tracks state in two places, the metastore (partitions) and the file system (files), so every time a job is planned it needs listing calls that grow with the number of partitions. On object storage those calls are slow, and in the past their results were even eventually consistent. Third, nobody agrees on what the table is. If someone accidentally drops one file into the directory, it becomes part of the table.

How it works — a four-layer tree

The first sentence of the spec overview changes direction: this format tracks individual files, not directories. A writer can put files anywhere and adds them to the table only through an explicit commit.

Layer Format What it holds
metadata.json JSON List of schemas, list of partition specs, properties, list of snapshots, current snapshot ID, refs
Manifest list Avro The manifests that make up one snapshot, each with its partition summary and file counts
Manifest Avro One line per data (or delete) file: path, partition values, row count, per-column lower and upper bounds
Data file Parquet, etc. The actual rows

A snapshot has a snapshot-id, a parent-snapshot-id pointing to its parent, a sequence-number that records commit order, and the path of its own manifest-list. The operation in the summary is one of append, replace, overwrite, or delete, and numbers such as the count of added rows and files are recorded alongside it. The data of a snapshot is the union of the live files listed in its manifests.

Each line (entry) in a manifest has a status — 0 EXISTING, 1 ADDED, 2 DELETED. Scan planning does not use DELETED entries. It then skips in two stages. It skips whole manifests using the partition summaries in the manifest list, and it skips files using the column statistics in the manifests. Directory listing appears nowhere.

What a commit does

Consider a second commit. It writes a new data file, writes a new manifest holding just that file, writes a new manifest list that points to both the first commit's manifest and the new manifest, and writes a new metadata.json with one more snapshot. The spec says manifests are reused between snapshots — old manifests are never rewritten, only pointed to. That is why the cost of a commit is proportional to the amount that changed, not to the size of the table.

Finally, the pointer in the catalog that says "the current metadata is this file" is swapped atomically. Anyone who opens the table after that moment sees the new state, and anyone who opened it before sees the old state all the way through. There is no half state. A sequence number grows by one with every commit, and the design means that a writer who lost a race and must commit again only needs to rewrite the manifest list.

What it looks like in the field

"The data I loaded yesterday is missing." When the files are in the bucket but the rows do not show up, it is almost always because those files are not in the manifests of the current snapshot — the write job wrote the files and died just before the commit. In Iceberg this is normal behavior. A file that was never committed is not part of the table.

"I deleted the directory and the table is fine," or the reverse. Data file paths are written in the manifests as full paths. The directory layout is only a convention and carries no meaning. If you delete a file by hand, the table breaks (a manifest points to a file that is gone), and if you drop a file in by hand, nothing happens.

Planning is slow. A table that was committed every minute by a stream ends up with thousands of manifests. Knowing that the cause of slow query planning can lie in the metadata layer, not in the data, is what leads you to compaction and manifest cleanup.

What really matters in practice

What you will do in the next lab

You create a table with Spark and commit one day of orders twice. Then you walk down the tree without any tools. You read the current and previous metadata paths from one row of the SQLite catalog, use jq to extract the snapshot list from metadata.json (parent, sequence number, manifest list), decode the Avro manifest list to confirm that the second snapshot points to the first commit's manifest unchanged, and decode the manifests to write down the live data files and row counts. The grader compares the values you wrote down with the actual metadata.