MiniMind — Train a Small Language Model Yourself, End to End
Measure Korean with the MiniMind tokenizer and train a small vocabulary yourself
Goal
Cut a Korean corpus with the vocabulary-6,400 tokenizer that MiniMind distributes and measure characters per token, then train a vocabulary of 1,024 with the same assembly (byte-level BPE) and compare. Record in numbers how the vocabulary size changes the compression ratio and the embedding share.
Why it matters
The token count is the amount of computation and the size of the context window. If the same text becomes twice the tokens, both training and inference become twice as expensive, and the text you can see at once is halved. But the vocabulary is one body with the model and cannot be changed once training has started — so only people who train from scratch make this decision. MiniMind's vocabulary was trained on Chinese and English data. To that vocabulary, Korean is almost a script it is seeing for the first time, so it gets split into bytes. This lab measures that cost and checks how short a vocabulary fitted to our corpus cuts. You also look at what share of the parameters the vocabulary takes in a small model — the reason MiniMind chose a small vocabulary of 6,400.
Steps
- Write the vocabulary size of the MiniMind tokenizer (
/opt/minimind/model/tokenizer.jsonandtokenizer_config.json), its bos, eos, and pad tokens, and their IDs to /root/mm/tok/minimind_vocab.json. - Cut the validation corpus (the text of
/opt/mm/data/pretrain_val.jsonl) with that tokenizer and write the tokens per character and per byte to /root/mm/tok/borrowed.json. - Train a byte-level BPE (vocabulary size 1024, with the special tokens
<|endoftext|>,<|im_start|>, and<|im_end|>in this order) on only the text of/opt/mm/data/pretrain.jsonland save it to /root/mm/tok/tokenizer.json. - Cut
<|im_start|>user\n안녕<|im_end|>\nwith the new tokenizer and write it to /root/mm/tok/chat_ids.json (the Korean word in the string means "hello"). - Take the same measurement as in step 2 with the new tokenizer and write it to /root/mm/tok/own.json.
- Write the characters per token for vocabularies of 384, 512, and 1024 to /root/mm/tok/sweep.json.
- Write the embedding share of three models — the MiniMind-3 default configuration, our small configuration, and our configuration with a vocabulary of 6400 — to /root/mm/tok/embed_share.json.
- In /root/mm/tok/report.md, write the three sections
## 빌려 온 토크나이저,## 우리 토크나이저, and## 어휘 크기와 임베딩(the Korean headings mean "Borrowed tokenizer", "Our tokenizer", and "Vocabulary size and embedding"), and include the characters per token from steps 2 and 5.
Notes
- If you type
mm-info, you see at a glance where the MiniMind copy, data, and reference outputs in the image are located. - Read a tokenizer with
from tokenizers import Tokenizer; Tokenizer.from_file(경로)(the placeholder is the file path), and useencode(글, add_special_tokens=False).idsto cut (the placeholder is the text). - Common mistakes: training without the special tokens so that
<|im_start|>is split into ten bytes, measuring the compression ratio on the training corpus instead of the validation corpus, and counting the shared embedding and output layer twice. - Sources: MiniMind train_tokenizer.py · HuggingFace tokenizers — BPE · ByteLevel pre-tokenization
Open MiniMind's vocabulary
Read /opt/minimind/model/tokenizer.json and tokenizer_config.json and write the vocabulary size, the bos, eos, and pad token strings, and their IDs to /root/mm/tok/minimind_vocab.json as vocab_size, bos, eos, pad, bos_id, eos_id, and pad_id.
The vocabulary size is Tokenizer.get_vocab_size(), and IDs come from token_to_id(). Which tokens are bos, eos, and pad is found not in tokenizer.json but in bos_token, eos_token, and pad_token of tokenizer_config.json.
Measure Korean with the borrowed vocabulary
Cut all the text in /opt/mm/data/pretrain_val.jsonl with the MiniMind tokenizer and write total characters ÷ tokens and total UTF-8 bytes ÷ tokens to /root/mm/tok/borrowed.json as tokens, chars_per_token, and bytes_per_token.
A Hangul character is 3 bytes in UTF-8. If characters per token is less than 1, it means one character was split into several pieces. If you make the measurement script take the tokenizer path and the result path as arguments, you can reuse it as it is in step 5.
Train a vocabulary of 1024
Write the script /root/mm/tok/train_tok.py (arguments: vocabulary size, save path), and using only the text of /opt/mm/data/pretrain.jsonl, train the same assembly as MiniMind (models.BPE + pre_tokenizers.ByteLevel(add_prefix_space=False) + decoders.ByteLevel) with vocab_size=1024 and the special tokens <|endoftext|>, <|im_start|>, and <|im_end|> in that order, and save it to /root/mm/tok/tokenizer.json.
It is BpeTrainer(vocab_size=…, initial_alphabet=pre_tokenizers.ByteLevel.alphabet(), special_tokens=[…]). The special tokens become IDs 0, 1, and 2 in the order written. After saving, it is normal for get_vocab_size() to come out smaller than 1024 — the pairs to merge ran out first.
Special tokens must be one piece
Cut the string <|im_start|>user\n안녕<|im_end|>\n with the new tokenizer and write it to /root/mm/tok/chat_ids.json as text (that string) and ids (the list of token IDs). The Korean word in the string means "hello".
Strings registered as special tokens are not split into bytes but become a single ID. The first ID must be 1 (<|im_start|>), and 2 (<|im_end|>) must appear once. SFT loss masking uses these IDs as markers.
Measure again with our vocabulary
Take the same measurement as in step 2 with the new tokenizer (/root/mm/tok/tokenizer.json) and write it to /root/mm/tok/own.json as tokens, chars_per_token, and bytes_per_token.
To compare the two numbers, it must be the same validation corpus and the same calculation. Our corpus has few kinds of words, so one word becomes one token even with a vocabulary of about 700 — with a real Korean corpus, a much longer tail would remain.
How much shorter does it get as you grow the vocabulary
Train vocabularies of 384 and 512 as well under the same conditions, and write the validation characters per token of the three vocabularies (384, 512, and 1024) to /root/mm/tok/sweep.json as {"384": 값, "512": 값, "1024": 값} (the placeholders are the measured values).
As the vocabulary grows, characters per token grows, but at some point it stops. If you recall how many entries the actual vocabulary had when you gave 1024 (step 3), you can see why it stops. The 1024 value must equal the value from step 5.
How much of the parameters does the vocabulary eat
Create three models with mmkit.new_model(설정) (the placeholder is the configuration) — the MiniMind-3 default (hidden_size, num_hidden_layers, and vocab_size of MiniMindConfig()), our small configuration (mmkit.SMALL), and our configuration with vocab_size 6400 — and write the embedding parameter count, total parameter count, and share to /root/mm/tok/embed_share.json as minimind3, ours, and ours_vocab6400.
MiniMind shares the embedding and the output layer (lm_head) (tie_word_embeddings). model.parameters() returns a shared tensor only once, so its sum is the total. The embedding is model.model.embed_tokens.weight.
Leave the reasons for choosing the vocabulary
In /root/mm/tok/report.md, write the three sections ## 빌려 온 토크나이저, ## 우리 토크나이저, and ## 어휘 크기와 임베딩 (the Korean headings mean "Borrowed tokenizer", "Our tokenizer", and "Vocabulary size and embedding"), and include the characters per token (chars_per_token) of step 2 and step 5 as numbers.
Write one line each on how many times longer the borrowed vocabulary cuts Korean, and what price that is in training and inference. In the last section, it is good to include what percent of the total the embedding becomes if you use a vocabulary of 6400 with our model size.