TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

Measure Korean with the MiniMind tokenizer and train a small vocabulary yourself

Continue in TT Lab

Goal

Cut a Korean corpus with the vocabulary-6,400 tokenizer that MiniMind distributes and measure characters per token, then train a vocabulary of 1,024 with the same assembly (byte-level BPE) and compare. Record in numbers how the vocabulary size changes the compression ratio and the embedding share.

Why it matters

The token count is the amount of computation and the size of the context window. If the same text becomes twice the tokens, both training and inference become twice as expensive, and the text you can see at once is halved. But the vocabulary is one body with the model and cannot be changed once training has started — so only people who train from scratch make this decision. MiniMind's vocabulary was trained on Chinese and English data. To that vocabulary, Korean is almost a script it is seeing for the first time, so it gets split into bytes. This lab measures that cost and checks how short a vocabulary fitted to our corpus cuts. You also look at what share of the parameters the vocabulary takes in a small model — the reason MiniMind chose a small vocabulary of 6,400.

Steps

  1. Write the vocabulary size of the MiniMind tokenizer (/opt/minimind/model/tokenizer.json and tokenizer_config.json), its bos, eos, and pad tokens, and their IDs to /root/mm/tok/minimind_vocab.json.
  2. Cut the validation corpus (the text of /opt/mm/data/pretrain_val.jsonl) with that tokenizer and write the tokens per character and per byte to /root/mm/tok/borrowed.json.
  3. Train a byte-level BPE (vocabulary size 1024, with the special tokens <|endoftext|>, <|im_start|>, and <|im_end|> in this order) on only the text of /opt/mm/data/pretrain.jsonl and save it to /root/mm/tok/tokenizer.json.
  4. Cut <|im_start|>user\n안녕<|im_end|>\n with the new tokenizer and write it to /root/mm/tok/chat_ids.json (the Korean word in the string means "hello").
  5. Take the same measurement as in step 2 with the new tokenizer and write it to /root/mm/tok/own.json.
  6. Write the characters per token for vocabularies of 384, 512, and 1024 to /root/mm/tok/sweep.json.
  7. Write the embedding share of three models — the MiniMind-3 default configuration, our small configuration, and our configuration with a vocabulary of 6400 — to /root/mm/tok/embed_share.json.
  8. In /root/mm/tok/report.md, write the three sections ## 빌려 온 토크나이저, ## 우리 토크나이저, and ## 어휘 크기와 임베딩 (the Korean headings mean "Borrowed tokenizer", "Our tokenizer", and "Vocabulary size and embedding"), and include the characters per token from steps 2 and 5.

Notes

Open MiniMind's vocabulary

Read /opt/minimind/model/tokenizer.json and tokenizer_config.json and write the vocabulary size, the bos, eos, and pad token strings, and their IDs to /root/mm/tok/minimind_vocab.json as vocab_size, bos, eos, pad, bos_id, eos_id, and pad_id.

The vocabulary size is Tokenizer.get_vocab_size(), and IDs come from token_to_id(). Which tokens are bos, eos, and pad is found not in tokenizer.json but in bos_token, eos_token, and pad_token of tokenizer_config.json.

Measure Korean with the borrowed vocabulary

Cut all the text in /opt/mm/data/pretrain_val.jsonl with the MiniMind tokenizer and write total characters ÷ tokens and total UTF-8 bytes ÷ tokens to /root/mm/tok/borrowed.json as tokens, chars_per_token, and bytes_per_token.

A Hangul character is 3 bytes in UTF-8. If characters per token is less than 1, it means one character was split into several pieces. If you make the measurement script take the tokenizer path and the result path as arguments, you can reuse it as it is in step 5.

Train a vocabulary of 1024

Write the script /root/mm/tok/train_tok.py (arguments: vocabulary size, save path), and using only the text of /opt/mm/data/pretrain.jsonl, train the same assembly as MiniMind (models.BPE + pre_tokenizers.ByteLevel(add_prefix_space=False) + decoders.ByteLevel) with vocab_size=1024 and the special tokens <|endoftext|>, <|im_start|>, and <|im_end|> in that order, and save it to /root/mm/tok/tokenizer.json.

It is BpeTrainer(vocab_size=…, initial_alphabet=pre_tokenizers.ByteLevel.alphabet(), special_tokens=[…]). The special tokens become IDs 0, 1, and 2 in the order written. After saving, it is normal for get_vocab_size() to come out smaller than 1024 — the pairs to merge ran out first.

Special tokens must be one piece

Cut the string <|im_start|>user\n안녕<|im_end|>\n with the new tokenizer and write it to /root/mm/tok/chat_ids.json as text (that string) and ids (the list of token IDs). The Korean word in the string means "hello".

Strings registered as special tokens are not split into bytes but become a single ID. The first ID must be 1 (<|im_start|>), and 2 (<|im_end|>) must appear once. SFT loss masking uses these IDs as markers.

Measure again with our vocabulary

Take the same measurement as in step 2 with the new tokenizer (/root/mm/tok/tokenizer.json) and write it to /root/mm/tok/own.json as tokens, chars_per_token, and bytes_per_token.

To compare the two numbers, it must be the same validation corpus and the same calculation. Our corpus has few kinds of words, so one word becomes one token even with a vocabulary of about 700 — with a real Korean corpus, a much longer tail would remain.

How much shorter does it get as you grow the vocabulary

Train vocabularies of 384 and 512 as well under the same conditions, and write the validation characters per token of the three vocabularies (384, 512, and 1024) to /root/mm/tok/sweep.json as {"384": 값, "512": 값, "1024": 값} (the placeholders are the measured values).

As the vocabulary grows, characters per token grows, but at some point it stops. If you recall how many entries the actual vocabulary had when you gave 1024 (step 3), you can see why it stops. The 1024 value must equal the value from step 5.

How much of the parameters does the vocabulary eat

Create three models with mmkit.new_model(설정) (the placeholder is the configuration) — the MiniMind-3 default (hidden_size, num_hidden_layers, and vocab_size of MiniMindConfig()), our small configuration (mmkit.SMALL), and our configuration with vocab_size 6400 — and write the embedding parameter count, total parameter count, and share to /root/mm/tok/embed_share.json as minimind3, ours, and ours_vocab6400.

MiniMind shares the embedding and the output layer (lm_head) (tie_word_embeddings). model.parameters() returns a shared tensor only once, so its sum is the total. The embedding is model.model.embed_tokens.weight.

Leave the reasons for choosing the vocabulary

In /root/mm/tok/report.md, write the three sections ## 빌려 온 토크나이저, ## 우리 토크나이저, and ## 어휘 크기와 임베딩 (the Korean headings mean "Borrowed tokenizer", "Our tokenizer", and "Vocabulary size and embedding"), and include the characters per token (chars_per_token) of step 2 and step 5 as numbers.

Write one line each on how many times longer the borrowed vocabulary cuts Korean, and what price that is in training and inference. In the last section, it is good to include what percent of the total the embedding becomes if you use a vocabulary of 6400 with our model size.