TT Lab
Get started
Learn Learning paths Courses

LLM Engineering

Tokenisers — What the Model Sees Is Not Letters

Continue in TT Lab

In one line

The model sees neither characters nor words but a sequence of token ids. The tokenizer is what defines that conversion, and the rules it uses determine cost, context limit and multilingual quality all at once. You cannot estimate token counts from character counts, so always count them with that model's tokenizer.

Why this was needed

Say a team gets its API bill at the end of the month. They had the same announcement summarized in an English version and a Korean version, and although the character counts are similar, the Korean version's input tokens are several times higher. The same day, a report comes in that only the Korean documents had their tails cut off because they hit the context limit. Both have the same cause: the model is looking at Korean sentences broken into smaller pieces. To understand why, you have to start with how a tokenizer is built.

The first attempt was word units. The problem is that the vocabulary is unbounded. Every word not in the dictionary becomes an unknown token and the information is lost. In a language like Korean, where particles and endings attach to words, the vocabulary explodes from variations of the same stem alone.

The second attempt was character units. The vocabulary gets small, but the sequences get too long. In a model with a limited context length, length is cost, and learning relationships between distant positions also becomes harder.

BPE (Byte Pair Encoding) is the answer found between the two. Starting from characters, it repeatedly merges the pair that occurs together most often into one. Frequently used words become a single token as a whole, and rare words stay as pieces. It fixes the vocabulary size while effectively eliminating unknown tokens.

How it works

The training procedure is surprisingly simple.

  1. Split the corpus into words and turn each word into a list of characters. Attach a special marker at the end so that word boundaries are not lost.
  2. Find the pair that is adjacent most often overall. If frequencies are equal, pick one with a fixed rule (for example, alphabetical order).
  3. Merge that pair into a single token and record this merge rule in order.
  4. Repeat steps 2 and 3 until you reach the vocabulary size you want. The final vocabulary size is the number of base characters plus the number of merges.

The tie rule looks trivial but it is decisive. If you break ties arbitrarily, the merge order differs from run to run even when training on the same corpus, and every rule and id after that drifts. A tokenizer is a component you can use only if it is reproducible.

Encoding applies the learned merge rules in exactly the recorded order. A small example shows why the order matters. Suppose the merge rules were learned in the order below.

merges (in order):  1) e r   2) w e
word:               l o w e r _
apply 1 then 2:     l o w er _
apply 2 then 1:     l o we r _

Two identical rules, yet changing only the order of application changes the resulting tokens. The fact that er was created first during training is itself the information "in this corpus, grouping as er was more common", so encoding has to follow that order to get the same split as in training. Decoding is easy by comparison: concatenate the tokens and turn the end-of-word marker back into a space.

The tokenizers of real models add a few things to this skeleton.

This explains the Korean cost problem. In UTF-8 a single Hangul syllable is 3 bytes (for example, 가 is EA B0 80). If a byte-level tokenizer did not see enough Korean in its training corpus, few merge rules that join a Hangul syllable into one are created, and in the worst case that syllable remains as three bytes, that is, three tokens. While an English word becomes one token as a whole, a single Korean syllable becomes several tokens. You cannot change the tokenizer either. If the vocabulary changes, the embedding matrix changes and the model has to be trained again.

What goes wrong in the field

Counting with a different tokenizer. When estimating cost or calculating a context budget, people often count the tokens of every model with one convenient library. Vocabularies differ from model to model, so the numbers are wrong, and the symptom shows up as "I calculated that it was within budget, but the output is cut off". Count tokens with the tokenizer of the model you will call, or record the usage the API returns in its response and calibrate against it.

Unicode normalization gets mixed in. Korean can express the same character in two ways. You can write it as one precomposed syllable (NFC), or as the initial, medial and final jamo concatenated (NFD). They look identical on screen, but the code point count is two to three times higher, the token count grows with it, and string comparison and search fail silently. Text that went through an older macOS file system as a file name, or came out of some document converters, looks like this. If "it is the same sentence but the token count is oddly high", suspect normalization first.

Special token strings get mixed into the input. If text pasted by a user contains the string of a special token such as an end-of-text marker, the question is whether to treat it as ordinary characters or as a control token. If it is interpreted as a control token, the input can behave as if it were cut off at that point. That is why tiktoken raises an error by default when it meets such input. Before you turn the error off, first understand why it is blocking.

Being surprised that the model cannot count letters. Token boundaries differ from character boundaries. When a model makes mistakes on "the third letter of this word" or on counting digits, it is not because the model is stupid but because it cannot see that unit in the first place. Do not leave such work to the model; handle it in code.

Whitespace and formatting eat tokens. Indentation, consecutive newlines and repeated markdown symbols take up more tokens than you would expect. When shortening a prompt, look at the token count, not the character count, and print the tokens yourself to see which part uses the most.

How to check

If you built or changed a tokenizer, always look at three things.

import unicodedata

text = open("corpus.txt", encoding="utf-8").read()
print(unicodedata.is_normalized("NFC", text))   # False means mixed forms
print(len(text), len(text.encode("utf-8")))     # characters vs bytes
# round trip: decode(encode(x)) must equal x for every line
# compression: characters / tokens, compare across languages

First, the round-trip check. The result of encoding and then decoding must not differ from the original by a single character. If you leave out the end-of-word marker, spaces cannot be restored, and that shows up right here. If even one line differs, compare that line with diff and find at which token it diverged.

Second, the compression ratio. This is the number of characters divided by the number of tokens. The larger the value, the more characters one token holds. If you measure the English and Korean versions of the same content side by side, you can explain with numbers why Korean costs more.

Third, read the beginning of the merge rules. The first few dozen merges are the most common pieces in the corpus. Check by eye whether the pieces you expected, such as particles and endings, appear, and whether odd pieces (whitespace characters, control characters) are mixed in. If you see odd pieces, the preprocessing is wrong.

What to read next

The reading right after this covers embeddings, which turn text into vectors, and then in the first lab you implement BPE from the ground up. You count character frequencies, build the base vocabulary, learn 120 merge rules while obeying the tie rule, encode the documents by applying those rules in recorded order, and then do the round-trip check that restores the original from the id sequence alone, and the compression ratio calculation. You use only Python, with no libraries.