TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

Separate what a small model can and cannot do with numbers

Continue in TT Lab

Goal

Measure the perplexity and BPB of the reference pretrained and SFT models, and build an overfitting curve with 100 documents. Measure questions seen and not seen in SFT, two-digit addition, positions outside the training length, and whether the pretrained model knows the facts, and leave in numbers what this model cannot do.

Why it matters

Even a model with low loss that answers all seen questions correctly collapses if it steps just one notch out of distribution. If you do not know that boundary, you will use the model in the wrong place. In a large model, several causes mix and look blurred, but in a 1-million-parameter model, "rules generalize and facts are memorized," "even what it knows cannot be brought out if the question shape differs," and "it forgets old abilities while learning a new format" each show up as sharp numbers. Evaluation is not only measuring what was newly taught. Only if you remeasure the original abilities after SFT and DPO is what was forgotten recorded — that is the first step of this lab.

Steps

  1. Write the total loss, perplexity, and BPB of the reference pretrained and SFT models over the whole validation corpus to /root/mm/eval/ppl.json.
  2. Train 300 steps on only the first 100 documents of the training corpus, leaving the training and validation loss every 25 steps in /root/mm/eval/overfit_log.csv and the weights in /root/mm/eval/overfit.pth.
  3. Write the reference SFT model's accuracy (ignoring speech style), split into questions seen in SFT and questions not seen, to /root/mm/eval/heldout.json.
  4. Write the answers and accuracy on 10 two-digit additions to /root/mm/eval/ood.json.
  5. In 256-token windows, write the mean loss at positions 1–127 and 128–255 to /root/mm/eval/length.json.
  6. Count whether the pretrained model can continue each village's specialty and guardian animal, and write it to /root/mm/eval/probe.json.
  7. In /root/mm/eval/report.md, write the three sections ## 무엇을 재나, ## 과적합, and ## 작은 모델이 못 하는 것 (the Korean headings mean "What is measured", "Overfitting", and "What a small model cannot do"), and include the pretrained model's BPB and the two-digit addition accuracy.

Notes

Perplexity and BPB

With all the windows made by cutting /opt/mm/ref/val.npy into non-overlapping 128-token windows, write the mean loss (weighted by the number of windows), perplexity (exp of the mean loss), and BPB of the reference pretrained (pretrain.pth) and SFT (sft.pth) models to /root/mm/eval/ppl.json as {"pretrain": {"loss", "ppl", "bpb"}, "sft": {…}}.

The denominator of BPB is the number of UTF-8 bytes you get by tok.decode-ing the tokens remaining after removing the first token of each window (ID 3 or higher, that is, excluding special tokens). Recall why the SFT model's perplexity is in the hundreds — what it learned at 1e-3.

Overfit on 100 documents

From /opt/mm/ref/train.npy, cut off only up to before the 101st bos (100 documents), create a new model with mmkit.seed_all(0), train 300 steps with batch 8, length 128, and lr 3e-3 (no schedule), and leave step,train_loss,val_loss every 25 steps (validation with mmkit.lm_loss(model, val, n_batches=4)) in /root/mm/eval/overfit_log.csv and the final weights in /root/mm/eval/overfit.pth.

Find the step where the validation loss was lowest. After that, the training loss keeps falling but the validation loss rises — this is the point where it began to memorize those 100 instead of the rules. The grader checks that the lowest point comes before the end, that it rose by 0.2 or more from the lowest to the end, and that the training and validation losses diverged by 0.5 or more.

Questions seen and not seen

With the reference SFT model, have it answer greedily the first 60 single-turn conversations from each of sft.jsonl (seen questions) and sft_val.jsonl (unseen questions), and write the accuracy after stripping the trailing 입니다 and 이다 (the Korean polite and plain endings), by question type (fact: questions containing the Korean word for "village"; add: addition), to /root/mm/eval/heldout.json as {"seen": {…}, "held_out": {…}}.

The unseen fact questions are ones that appeared in the pretraining corpus but whose question shape was never seen in SFT. Compare them with the accuracy on unseen addition pairs — rules and facts generalize differently.

Two-digit addition

Ask the reference SFT model 10 two-digit additions (12+15, 23+41, 30+30, 45+12, 17+21, 50+25, 11+11, 34+52, 26+13, 40+19) with '{a} 더하기 {b}는?' (the Korean phrase means "what is a plus b?"), and write the share of answers that start with the correct number to /root/mm/eval/ood.json as accuracy and rows (q: in the form "12+15", answer: the model's answer).

The model saw only 100 one-digit additions. It keeps the answer format (a number plus the polite or plain ending), but the number comes out like the answer to a one-digit addition — it did not learn the "range" of the rule.

Positions outside the training length

Feed all the windows made by cutting val.npy into 256 tokens into the reference pretrained model, split the next-token loss per position into positions 1–127 (loss_pos_1_127) and 128–255 (loss_pos_128_255), average each, and write it to /root/mm/eval/length.json together with windows.

Get the per-position loss with cross_entropy(…, reduction="none") and split it into two stretches. This model was trained only on 128-token windows. RoPE uses relative position, so it does not collapse, but it is less accurate at far distances it has never seen.

Does the pretrained model know the facts

For each village in /opt/mm/data/world.json, feed the reference pretrained model [bos] + '{마을} 마을의 특산물은' and '… 수호 동물은' (the Korean phrases mean "the specialty of {village} village is" and "... the guardian animal is"), continue for 6 tokens greedily, count how many start with the correct answer, and write it to /root/mm/eval/probe.json as specialty_known, animal_known, and held_out_known_by_pretrain (the number correct among the 4 held-out facts).

Here you separate whether the SFT model got the unseen fact questions wrong because "it does not know" or because "it knows but cannot bring it out in question form." The guardian animals cycle through six in village order, so they are easy to memorize, while the specialties are paired one by one across twelve.

What this model cannot do

In /root/mm/eval/report.md, write the three sections ## 무엇을 재나, ## 과적합, and ## 작은 모델이 못 하는 것 (the Korean headings mean "What is measured", "Overfitting", and "What a small model cannot do"), and include the pretrained model's bpb from step 1 and the accuracy from step 4 as numbers.

In the last section, write one line each on what it "cannot do," by cause (facts it does not know, facts it knows but cannot bring out, ranges it has never seen, and forgotten abilities). Where you must not use this model is exactly this list.