Transformers — Compute Attention By Hand
The Same Sentence, a Different Token Count
In one line
A token is neither a character nor a word. It is a piece fixed when the vocabulary is built, and because the vocabulary differs from model to model, the token count for the same sentence differs from model to model.
Why this was needed
Both the price list and the context limit are in tokens. But what you can count on screen is only characters, so for a while you estimate with "multiply the character count by something, I guess". The question is when that estimate breaks.
It always breaks in the same places. You put in a Korean document and it costs far more than expected. The same content in English is cheap. You put in a log as it is, and the tokens explode even though the text is short. Only with user input mixed with emoji does the length calculation go off.
All of these have the same cause. A token is not a character unit. Which pieces count as one is written in the vocabulary, and that vocabulary was built by looking at training data. If the vocabulary was built from data with a lot of English, common English pieces are one token and Korean is chopped into small parts.
What is worse is that this difference is silent. No error occurs. It only shows up in the bill and the context limit.
Why start from bytes
The old approach kept a word dictionary and lumped every word not in the dictionary into a single <UNK>. Names not in the dictionary, typos, new slang and emoji all became the same token, and there was no way to reverse it.
The approach in use now starts from bytes. When you write text in UTF-8, any character becomes 1 to 4 bytes, and a byte has only 256 possible values, from 0 to 255. If you make those 256 the initial vocabulary, the notion of a character outside the vocabulary never arises.
"A".encode("utf-8") # b'A' → 1바이트
"가".encode("utf-8") # b'\xea\xb0\x80' → 3바이트
Here the cost of Korean shows up plainly. A single Hangul syllable is 3 bytes in UTF-8. If the vocabulary is small and has few rules for merging, one Hangul character becomes three tokens. An English letter is 1 byte, so under the same conditions one character is one token. That means the starting line differs by a factor of three.
How the vocabulary is built
What BPE (Byte Pair Encoding) does can be written in one sentence — merge the two items that most often appear together into one, and repeat until the vocabulary reaches the size you want.
One round takes three steps.
- Count how many times each pair of neighbors appears in the current list.
- Pick the pair that appears most often.
- Replace every place that pair occurs with one new number. New numbers go up one at a time from 256.
ids = [104, 101, 108, 108, 111] # "hello"
counts = {(104,101):1, (101,108):1, (108,108):1, (108,111):1}
# 전부 1회라 동점이다 — 동점을 어떻게 깰지 정해 두지 않으면
# 같은 글로 돌려도 어휘가 매번 달라진다.
Tie handling looks trivial but matters in practice. If training is not deterministic, the vocabulary you made yesterday and the one you make today differ, and then the model today reads data you encoded yesterday differently. Hugging Face's explanation of BPE also pins down the rule at the same point.
The point where you stop is the vocabulary size. If you stop at 256, there are no merges at all, and if you stop at 50,000, there are about 49,000 merge rules. The vocabulary size is a value set when the model is built, and if you change it, the size of the embedding table changes with it.
However, the vocabulary size is only an upper bound. When there are no more pairs worth merging, it stops there. Merging a pair that appears only once uses up one vocabulary slot and does not reduce the tokens at all. With little data, even if you ask for 50,000 it ends at a few thousand — raising the vocabulary size and the vocabulary actually growing are different things.
Encoding follows the learned order
When you encode new text after learning all the rules, you must keep the order in which they were learned.
The reason is that later rules use the results of earlier rules as their material. The rule that made number 256 must run first so that the rule for 257 can see 256. If you shuffle the order, the same rule list produces different tokens, and a model trained that way reads different text in service.
Decoding is the reverse. You expand a new number into two, and if one of those two is again a new number, you expand it again. When nothing is left, only bytes remain, and reading them as UTF-8 gives the original. It must not differ from the original by a single character — if this round trip breaks, no one can trace back what the model read.
What it looks like in the field
First, the bill for a Korean service comes out two or three times what was expected. The estimate applied the tokens-per-character figure guessed from English as it was. There is only one fix — measure with the actual sentences of your own service.
Second, you hit the context limit without warning. By character count there is plenty of room, but by tokens you are already over. If the code that cuts long documents works by characters, it cuts too much in some languages and cannot cut in others.
Third, the content is the same, but just changing the format increases the tokens. JSON with deep indentation, logs with many line breaks, and documents with tables aligned by spaces are like that. Spaces and line breaks are bytes too, and if the vocabulary lacks that piece, each one becomes a token.
Fourth, with emoji and rare characters only the length calculation goes off, and no error occurs. It is at the byte level, so it does not break. But a single emoji is 4 bytes, and if it is not in the vocabulary it eats three or four tokens.
Fifth, if you change the tokenizer, you have to retrain the model. The tokenizer is not a preprocessing tool outside the model; it is part of the model. If the numbers change, the rows of the embedding table change, and that is a different model.
What really matters in practice
- Do not estimate by character count; measure with your own sentences. When language and format change, the ratio changes too.
- Vocabulary size is not free. Raising it reduces tokens but enlarges the embedding table, and the amount of reduction keeps getting smaller. You should decide by looking at that point in numbers.
- Training must be deterministic. If you do not pin down even the rule for breaking ties, the vocabulary cannot be reproduced.
- Pin the round trip down with a test. Encoding, decoding, and checking that it equals the original is a one-line test, but if it breaks, everything below loses its meaning.
What you will do in the next lab
You grow /root/work/tf-token/bpe.py one step at a time. Instead of calling the tokenizer of a real model, you build the same algorithm yourself with only the standard library — this Pod has no transformers, no tokenizers and no tiktoken, and even numpy exists only inside /opt/onnx-lab/bin/python. So every number that comes out here is measured with a vocabulary you built yourself.
You start by opening with UTF-8 bytes and build up through counting adjacent pairs, merging one pair, learning merge rules, encoding in the learned order, and exact decoding. Then you change the vocabulary size and measure the curve along which the token count for the same text decreases.
The last step is the point of this lab. With a Korean paragraph and an English paragraph of the same content, you learn two vocabularies — one learned by looking at both texts together, the other learned by looking at English only. The sizes are the same. Then you encode the same paragraph with both vocabularies and set the token counts side by side. The vocabulary learned from English only has no rule at all that merges Hangul bytes, so Korean barely shrinks from its byte count. Even so, decoding is exact. You will see in numbers that the loss and the robustness come from the same structure. The grader actually imports your module, pokes at the functions directly with different inputs every time, and checks them against values it computes separately.