TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

A borrowed tokenizer splits Korean into bytes

Continue in TT Lab

In one line

A tokenizer is a dictionary that turns text into the numbers (tokens) a model consumes. MiniMind uses a byte-level BPE with a vocabulary of 6,400, a size chosen so that embeddings do not eat up the parameters of a small model. But that vocabulary was trained on Chinese and English, so it cuts Korean almost byte by byte. In this module, you train a new vocabulary on our corpus and compare, in numbers, how many tokens the same text becomes.

Why this was needed

A language model does not know characters. It receives only a line of integer IDs, and for each ID it pulls out one row of an embedding vector. So two things are decided at once.

MiniMind also recommends that you not retrain the tokenizer. When the vocabulary changes, the weights, data format, and inference interface all change with it, so you can no longer share models with others, and perplexity (PPL), which is computed per token, cannot be compared across models if the vocabulary differs. This course knowingly retrains it anyway — to measure for ourselves how expensive a borrowed vocabulary is on a Korean corpus.

How it works

MiniMind's trainer/train_tokenizer.py assembles three parts with HuggingFace tokenizers.

tok = Tokenizer(models.BPE())
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=False)
trainer = trainers.BpeTrainer(vocab_size=6400,
    initial_alphabet=pre_tokenizers.ByteLevel.alphabet(),   # 바이트 256개를 처음부터
    special_tokens=["<|endoftext|>", "<|im_start|>", "<|im_end|>", ...])
tok.decoder = decoders.ByteLevel()

Byte-level (ByteLevel) first turns text into UTF-8 bytes and takes the 256 bytes as the base characters. Any character can be represented as bytes, so there is no "unknown character". A Hangul character is 3 bytes in UTF-8, so if the vocabulary has no token that combines that character, one character becomes up to three tokens.

BPE repeats merging the two pieces that most often appear together into one until the vocabulary is full. If the pairs to merge run out first, it stops without filling the vocabulary size — the corpus of this course has few kinds of words, so even if you give 1,024, it stops at about 700. You will see this yourself in the lab.

Special tokens are placed in the very first slots before training. MiniMind uses ID 0 <|endoftext|> (padding), ID 1 <|im_start|> (bos), and ID 2 <|im_end|> (eos), and it reserves 36 slots, including tokens for tool calls and thinking mode. These IDs are markers that the chat format and loss masking rely on, so they must not be split into bytes.

What it looks like in the field

If you want to fine-tune an open model on internal documents and the context window fills quickly so long documents get cut off, first measure how many tokens that model's tokenizer cuts Korean into. The amount of text you can fit at the same cost varies two to three times depending on the tokenizer. However, you cannot change the vocabulary of an already trained model — the moment you change the vocabulary, the embedding table becomes useless. Choosing a vocabulary is a decision you can make only when training from scratch, and that is why MiniMind keeps using its vocabulary once it is set.

When comparing quality between models, you must not compare per-token loss as it is if the vocabularies differ. If tokens are long, hitting one token is harder, and if they are short, it is easier. This is why the MiniMind README recommends using bits per byte (BPB) when the tokenizers differ, and in the last module of this course you measure BPB yourself.

How this course differs from the original MiniMind

MiniMind's train_tokenizer.py trains a vocabulary of 6,400 on the conversation content of the SFT data (sft_t2t_mini.jsonl) and reserves 36 slots for special tokens — including tool calls (<tool_call>), thinking mode (<think>), and spare slots for future use (<|buffer1|> …). This course changes three things. It uses only the text of the pretraining corpus as training data (so that the grader can retrain under the same conditions and compare), it reduces the vocabulary to 1,024, and it keeps only the three special tokens that chat really needs. This is because it does not deal with tool calls or thinking mode. Instead, it keeps the same names and the same IDs (0, 1, 2) as MiniMind, so that the chat format and loss masking in later modules behave exactly like the MiniMind code.

What you will do in the next lab

You cut our Korean corpus with MiniMind's vocabulary and measure characters per token, then train a vocabulary of 1,024 with the same assembly and compare. You also record in numbers how the compression ratio changes as you vary the vocabulary among 384, 512, and 1,024, and how much the vocabulary size changes the embedding share.