TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

In a small model the line between memorizing and learning is sharp

Continue in TT Lab

In one line

Evaluation goes beyond "the loss is low" to separating what the model can do and what it cannot. You measure quality as a language model with perplexity and BPB, see overfitting from the gap between training and validation loss, and record in numbers how it collapses on questions, digit counts, and lengths it has never seen. In a 1-million-parameter model those boundaries are very sharp, so you can pull apart one by one the phenomena that look blurred together in a large model.

Why this was needed

In the earlier modules, the loss went down well, and the SFT model answers all the questions it saw in training correctly. But can you call this model a "chatbot that knows about the villages"? The MiniMind README includes conversation examples of its own zero model as they are — it answers Chinese questions plausibly but strings together meaningless words for English questions. And it writes that "factual knowledge and generalization ability are still limited." Evaluation is the work of finding where that "limit" begins, and if you do not know that boundary, you will use the model in the wrong place.

How it works

Perplexity and BPB. Perplexity is exp(mean loss), read roughly as "how many candidates the model hesitates between per token on average." But it depends on the size of the tokens. BPB is the total loss (nats) converted to bits by dividing by ln 2 and divided by the byte count of the same text, so it can be compared even when tokenizers differ. Our reference pretrained model has a validation perplexity of 2.48 and BPB of 0.21. At the same point, the reference SFT model's perplexity is in the hundreds — it learned only the chat format at the large learning rate of 1e-3 and forgot the ability to continue the corpus (catastrophic forgetting).

Overfitting. If you train for a long time on 100 documents, the training loss keeps falling but the validation loss starts to rise from some point. That is the point where the model began to memorize those 100 instead of the rules. The corpus of this course consists of sentences stamped out from a template, so overfitting comes late — with real text, the gap would open earlier.

What it has seen and what it has not. The reference SFT model answers 100% of the questions it saw in SFT. But for facts that were in the pretraining corpus but for which it never saw that question shape in SFT, it is 0%. To separate the reasons, you must first see whether the pretrained model knows that fact — if you have it continue after "the specialty of village ○○ is," it gets most of the guardian animals right but only a few of the specialties. It cannot bring out even what it knows in the form of a question, and of course cannot bring out what it does not know. Addition is different — it gets 77% of one-digit addition pairs it has never seen. Repeating rules generalize, while facts that must be memorized one by one do not.

Out of distribution. If you ask a model that has seen only one-digit addition about two-digit addition, it gives a one-digit answer in the right format. If you feed 256 tokens into a model trained only on 128 tokens, the loss at the later positions goes up. RoPE uses relative position, so it does not collapse, but it is less accurate at distances it has never seen. MiniMind handles this problem with YaRN (inference_rope_scaling), which stretches RoPE at inference, but the argument description in eval_llm.py says it solves only the position encoding problem — it does not give the ability to handle long text itself.

What it looks like in the field

When evaluating an internal model, if you gather only questions similar to those used in training, the score looks inflated. You need a separate evaluation set that moves the question shape, topic, and length one step at a time outside the training distribution for the boundary to show. Also, after SFT and DPO you must remeasure the original abilities (perplexity on corpus continuation, general knowledge) — if you measure only what was newly taught, what was forgotten is recorded nowhere. When you look at wrong answers, first separate "does it not know, does it know but fail to bring it out, or is it out of range." The three have different remedies — if it does not know, add data; if it fails to bring it out, add SFT data with varied question shapes; and if it is out of range, put in data for that range or narrow the use.

How this course differs from the original MiniMind

MiniMind's eval_llm.py is a tool that loads the trained weights, has them answer a few prepared questions, and prints speed (tokens/s) — a human reads it and judges. Separately, MiniMind measures the model with external benchmarks such as C-Eval, C-MMLU, and OpenBookQA — with tens of thousands of multiple-choice items, it can tell whether even a 64M model is better than random. A 1M model cannot be distinguished from random on such benchmarks, so this course closes the world (12 villages and 100 additions) and measures exactly by splitting, within it, what was seen, what was not seen, and what is out of range. The advantage of a closed world is that you can separate the reasons for wrong answers one by one, and the disadvantage is that you cannot carry the numbers over to the outside world. This model's 100% accuracy only means "it solves the village quiz well," not that it understands language. When you build an evaluation set, you must always first write down what that set represents.

What you will do in the next lab

You measure the perplexity and BPB of the two reference models and build an overfitting curve with 100 documents. You measure in turn the accuracy on questions seen and not seen in SFT, two-digit addition, the loss outside the training length, and whether the pretrained model knows the facts, and leave what this model cannot do in a report.