TT Lab
Get started
Learn Learning paths Courses

Data Pipelines

How Files Are Laid Out Sets the Cost: Partitions and Compaction

Continue in TT Lab

Goal

You build a tool pq.py that handles a partition lake on a filesystem. You measure the file count and size distribution while changing the key, measure how much a condition on the partition key reduces the number of files opened, run compaction, check what a reader sees in the middle of compaction, and measure the cost of changing the key.

Why it matters

Even if you insert the same rows, the number of files a query opens differs by tens of times depending on how the files are laid out. Partitioning divides directories by value, and a file you do not open costs 0. But that benefit arises only when you filter by a field that is in the directory name. That said, if you put every frequently used field in as a key, the number of partitions grows by the product of the values and a single file becomes a few lines. The fixed cost of opening a file becomes greater than the cost of reading its contents, the number of entries in the listing explodes, and compression does not work. So choosing a key is always a trade-off. When small files pile up, you run compaction. What is hard is not the compaction itself but what a reader sees in the middle of it. A reader that collects files by scanning the directory picks up both old and new files and counts the same lines twice. If you make the reader look at the listing and do the listing swap only once at the end, that in-between disappears. We do not use Parquet. The lab image has no pyarrow and the Pod cannot install at runtime. Concepts such as row groups were covered in the reading through the official documentation, and here we build the same structure by hand with partition directories, a manifest, and JSON Lines. The grader does not trust the text you write down. It sets up the raw data that the grader created in a temporary work folder and actually runs your tool, and checks that the size written in the manifest matches the actual size on disk, that lines are preserved, and that the number of files cut away is correct. The raw data and the target size change on every run.

Steps

  1. Create and run /root/parts/gen_orders.py to create /root/parts/work/raw.jsonl.
  2. In /root/parts/pq.py, create write so that it builds a lake split by one key and a manifest.
  3. Add layout to produce the file count, the size distribution, and the number of small files.
  4. Make write accept several keys and be able to limit the number of lines in a file, and build a finely split lake.
  5. Add query to cut away the files to open using a condition on the partition key.
  6. Add compact to combine small files within one partition to be close to the target size.
  7. Add --crash=before-swap and --via=glob and check what a reader sees in the middle of compaction.
  8. Add repartition to change the key and rewrite, and write down the cost in /root/parts/work/partition_report.md.

Reference

Create raw data with several partition key candidates

Create and run /root/parts/gen_orders.py to create /root/parts/work/raw.jsonl. There must be at least 200 lines, one line holds order_id, day, region, channel, and an integer amount, and there must be at least 4 kinds of day, at least 3 kinds of region, and at least 3 kinds of channel.

The number of distinct values matters. This is because what you will see in this lab is how many partitions you get when you take all three as keys. If you multiply the number of distinct values and divide by the number of raw lines, you can see in advance how many lines will be left per file. You must fix the seed so that the raw data does not wobble while you compare changing the keys.

Split by one key and leave a listing

In /root/parts/pq.py, create write <작업폴더> --lake=<이름> --key=<칸> (the placeholders stand for the work folder, the lake name, and the field) so that it makes a 칸=값 directory (field=value) for each value, writes part-0000.jsonl under it, and leaves _manifest.json inside the lake.

The manifest's bytes must be the actual file size on disk. If you write a value computed by adding up line lengths, it is off because of newlines or encoding, and that discrepancy later cuts at the wrong place in compaction. Measure the size again after you write the file. If you also write the manifest under a temporary name and swap it in, the listing is never read half written.

Measure how the files are laid out by size

Add layout <작업폴더> --lake=<이름> --small=<바이트> (the placeholders stand for the work folder, the lake name, and bytes) to produce the file count, partition count, line count, total bytes, the average, median, minimum, and maximum size, and the number of files below --small.

The average lies. If one big file and hundreds of small files are mixed, the average looks fine. Take the median by the nearest rank and do not interpolate. avg_bytes is the quotient (rounded down) of the total bytes divided by the number of files. Compute all of these values by reading from the manifest.

Split finely to make small files

Make write accept several keys, as in --key=day,region,channel, and be able to limit the maximum number of lines in one file with --rows=<줄 수> (the placeholder stands for the number of lines). Then build a finely split lake and compare it with the earlier lake using layout.

Every time you add one key, the partition count is multiplied by the number of distinct values of that field. Stack the directories nested in key order — day=2026-01-03/region=seoul/channel=app/. If --rows is absent or 0, it is one file per partition. If you put the small_files of the two lakes side by side, you can see in numbers what was lost.

A file you do not open costs 0

Add query <작업폴더> --lake=<이름> --where=<칸=값[,칸=값]> (the placeholders stand for the work folder, the lake name, and field=value pairs) to produce the number of matching lines and the amount, but cut away the files to open only by a field that is in the partition key. The response includes both files_total and the files_scanned actually opened.

A field not in the key is not in the directory name, so you cannot cut anything by that condition. In that case files_scanned must equal files_total. Counting this honestly is the whole point of this step — if you inflate it here, you will never find later what is slow. Compare values as strings.

Combine small files

Add compact <작업폴더> --lake=<이름> --target=<바이트> (the placeholders stand for the work folder, the lake name, and bytes) to concatenate the files within one partition in path order, but cut when adding the next file would exceed the target. After writing all the new files, swap the manifest, and then delete the old files.

The order matters. If you swap the manifest first, the reader opens a file that does not exist yet. The new file names must not overlap the old names — if they overlap, you overwrite the file you are reading and that partition becomes entirely empty. If you put a generation number in the name, they never overlap. The line count must be the same before and after compaction.

What a reader sees in the middle of compaction

Add --crash=before-swap to compact so that it dies with exit code 9 after writing all the new files, just before swapping the manifest, and add --via=glob to query to build a read that ignores the manifest and scans part-*.jsonl in the directory. After killing it, compare the answers of the two reads.

This is the key scene of this lab. The reader using the manifest sees no change, and the reader scanning the directory counts the same lines twice. Even after it dies, the old files and the old manifest must remain as they were. Then, when you finish compaction properly, the old files are deleted and the two reads become the same again.

Put a price on changing the key

Add repartition <작업폴더> --from=<이름> --to=<이름> --key=<칸[,칸]> (the placeholders stand for the work folder, the lake names, and the field names) to read everything following the manifest of the --from lake, not the raw files, and rewrite it with the new key. Then write in /root/parts/work/partition_report.md four sections: ## 무엇을 어떻게 쪼갰나 ## 작은 파일 문제 ## 묶기 ## 파티션을 바꾸는 비용 (in order: what was split and how, the small file problem, compaction, and the cost of changing the partition).

The partition key is the directory structure itself, so changing it means rewriting every row. Whether rows_read and rows_written are equal is the first check, and bytes_read and bytes_written are the actual values of that work. In the report, write those byte counts as numbers — they are the only numbers you need when someone proposes changing the key in the next meeting.