TT Lab
Get started
Learn Learning paths Courses

LLM Engineering

Chunking — Deciding the Retrieval Unit Is Half the Job

Continue in TT Lab

In one line

A chunk is the smallest unit of retrieval and also the unit that goes into the context. If it is too big, noise gets mixed in; if it is too small, the context is cut. There is one criterion for a good chunk: you should be able to answer one question from reading that chunk alone.

Why this was needed

Say you indexed a 40-page operations manual whole, as a single item. If you ask "what is the backup retention period?", the manual comes out first. Retrieval succeeded. But putting 40 pages in the context exceeds the token budget, and even if you could, the one line with the answer is buried in the middle of a long text and the model misses it. Conversely, if you cut the manual line by line and index it, the line "that value is 30 days" comes up, but what the value belongs to is in the previous line and cannot be known.

If you index documents whole, two things collapse. First, a long document holds several topics, so its vector is smeared into an average. Second, even when retrieval succeeds, you have to put the whole document into the context, so tokens explode. So you have to cut documents, and how you cut them decides retrieval quality. And once you pick a method, it is expensive to change. If you change the chunking strategy, you have to re-index everything.

How it works

Fixed-length splitting cuts every N characters or N tokens. It is simple to implement and the sizes are predictable, but it breaks in the middle of sentences and destroys meaning. When measuring length, tokens are better than characters, because the limits of both the embedding model and the generation model are set in tokens.

Delimiter-based splitting cuts at paragraph or sentence boundaries. It preserves units of meaning, but lengths become uneven. In practice you mix the two methods: split by paragraph first, re-split only the paragraphs that are too long into sentences, and merge fragments that are too short with their neighbors.

Overlap is a technique that lets adjacent chunks share part of their content. It eases the problem where information that straddles a boundary is not retrieved from either side. The price is storage and duplicate retrieval. If you overlap 50 tokens on 400-token chunks, a chunk starts every 350 tokens, so the number of chunks rises by about 14%, and the same content can appear twice in the top results, wasting space in the context.

Hierarchical splitting searches with small chunks and puts the larger unit they belong to (a paragraph or a section) into the context. It is a compromise that tries to get retrieval precision and context completeness together.

Metadata fills in the missing context of a chunk. If you attach the document title, section path and update date to a chunk, it adds context to search results and can also be used for filtering. Just prepending the title and section name to the chunk text when indexing often raises retrieval quality noticeably.

Going back to a practical criterion for choosing size: if, when you pull out a single chunk and read it, you cannot tell what "it" refers to, the context is insufficient, and if one chunk holds three unrelated topics, the vector has been smeared.

What goes wrong in the field

Cutting at periods cuts the wrong places. A period does not appear only at the end of a sentence. Sentences get split at 3.5, v1.2, domain names, abbreviations and ellipses. Conversely, parts without periods, such as lists, tables and code, stay joined and become giant chunks. The symptom is that at both ends of the chunk length distribution you see pieces of a few characters together with lumps of thousands of characters.

Tables and code blocks get torn. With fixed-length cutting, the header and body of a table are separated and both become useless. A chunk with only numbers cannot tell you what the numbers are. You need special handling: repeat the header in every piece, or keep a table whole as one chunk.

The tail of a chunk is silently cut. An embedding model has an input length limit, and text longer than that is usually cut, keeping only the front. For example, sentence-transformers uses input that exceeds the model's max_seq_length only up to that length, and a common value in the BERT family is 512 tokens. No error appears. The content in the latter part of the chunk is not reflected in the vector and is never retrieved. Set the maximum chunk length smaller than the embedding model's limit.

Chunk numbers shift. If you number chunk ids only as "the Nth piece of the document", adding just one sentence at the start of a document shifts every later number by one. The correct-answer labels of the evaluation set, caches and user-feedback records all end up pointing to the wrong chunks. For a system meant to last, make a stable id from the document id and a hash of the content.

How to check

Once you have made the chunks, look at the distribution and samples before indexing.

import statistics, random

rows = [line.rstrip("\n").split("\t", 2)
        for line in open("chunks.tsv", encoding="utf-8")]
lens = [len(r[2]) for r in rows]
print(len(rows), min(lens), statistics.median(lens), max(lens))
print(sum(1 for n in lens if n < 15), "very short chunks")
for r in random.sample(rows, 5):
    print(r[0], r[1], r[2])

Among the numbers, look at the minimum and maximum first. If there are many very short pieces, the splitting rule is cutting at places that are not sentence boundaries, and if there are very long pieces, there are regions without delimiters. Read five random samples aloud. If you cannot say "what question can I answer with this" from a single chunk, that chunk is useless even when it is hit by search.

If you have evaluation questions, do one more thing: check whether each question's answer sentence is contained whole in a single chunk. If an answer spans two chunks, that question cannot be fully answered by any retriever. Finally, if you run the same queries against a document-level index and a chunk-level index and compare Recall@K, you can say in numbers whether chunking actually helped.

What you will do in the next lab

You cut 30 documents by period into about 90 chunks and save them, following the rule of stripping surrounding whitespace and discarding empty pieces. You build a TF-IDF index at that chunk level and search 8 queries, measure retrieval quality with Recall@3 and MRR, and then complete a full RAG pipeline, including context assembly, grounding judgment and refusal of questions outside the corpus. You also see for yourself where splitting at periods cuts awkwardly.