TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

SFT teaches only the answer — loss masking draws that line

Continue in TT Lab

In one line

A pretrained model only continues a corpus and does not answer questions. SFT (supervised fine-tuning) turns question–answer conversations into one line in the chat format and trains with the same next-token prediction, but computes the loss only on answer (assistant) tokens. MiniMind's generate_labels draws that boundary. In this module, you train a masked model and an unmasked model for the same number of steps, and see in numbers how the loss on question tokens and the loss on answer tokens diverge.

Why this was needed

To teach a model that "when it gets a question, it answers," you need two things. One is a marker the model can recognize for where the question ends and where the answer begins, and the other is the choice of what to learn. If you train on the whole conversation, the model learns even how to write the question — it spends much of its learning signal imitating what a user might say, and after the answer ends it may even make up the next question itself. What we want is a model that treats the question as a condition and produces the answer.

How it works

Chat format. MiniMind uses the same ChatML shape as the Qwen family.

<|im_start|>user
가람 마을의 특산물은 뭐야?<|im_end|>
<|im_start|>assistant
가람 마을의 특산물은 인삼입니다.<|im_end|>

MiniMind's actual template inserts an empty <think>\n\n</think>\n\n at each assistant turn and removes it with 80% probability during training (a device to fit thinking mode). With 20% probability it also attaches a system prompt at the front. This course does not deal with thinking mode, so it drops both and uses only the shape above (mmkit.chat_text).

Label masking. generate_labels finds the token pieces of <|im_start|>assistant\n in the token sequence, and from there until the <|im_end|>\n pieces end, it puts the original tokens in the labels. Everything else is -100. The model's cross_entropy(..., ignore_index=-100) skips those slots.

labels = [-100] * len(input_ids)
# <|im_start|>assistant\n 을 찾으면 그 뒤부터 <|im_end|>\n 까지 labels[j] = input_ids[j]

It matters to include the end marker <|im_end|> in the correct answer. That is how the model learns when to stop the answer. In a multi-turn conversation, this stretch occurs at every assistant turn. Counting on our SFT data, only a little over 22% of the real tokens go into the loss — the other 78% are only conditions.

Hyperparameters. MiniMind's train_full_sft.py starts from the pretrained weights and runs two epochs at lr 1e-5 (one fiftieth of pretraining's 5e-4). This choice is meant not to shake the language ability already learned. The small model in this course sets a much larger 1e-3, and at that price it forgets quite a lot of the ability to continue the pretraining corpus — you measure it with perplexity in the last module.

What it looks like in the field

If you did SFT on internal support records and the model continues after an answer with a fake question that starts with "Customer: …", masking was missing or you left the end marker out of the correct answer. Conversely, if you do not include the end marker, the model does not stop and writes to the maximum length. Both symptoms show up right away if you print one line of labels — the debugging output left as a comment in MiniMind's SFTDataset.__getitem__ (printing the input tokens, next tokens, and labels side by side) is for exactly that purpose.

Another common misconception is the expectation that "SFT puts knowledge in." What SFT mainly teaches is format and attitude — in what shape, and where to stop, when it answers a question — and facts mostly come from pretraining. If you try to put in facts that are not in the data with a few thousand SFT examples, the model only memorizes those sentences and cannot bring them out when asked in a different question shape. In the last module of this course, you see that phenomenon in numbers.

How this course differs from the original MiniMind

MiniMind's SFT data is mixed with tool-call conversations (the system turn holds the tool list, the assistant calls with <tool_call>, and the tool role returns the result), and pre_processing_chat does not touch conversations that have tools. The template wraps tool results as a user turn, so those slots also drop out of the labels — the model does not imitate tool results, and learns only how to continue answering after seeing them. This course does not deal with tool calls and uses only one- or two-turn question–answer pairs and greetings. The maximum length is also 64 rather than MiniMind's recommended value (768) — even our longest conversation is about 40 tokens, so at 64 nothing is cut and there is little padding. Conversely, with real data, you must first measure the conversation length distribution and then decide the maximum length. If it is too short, the end of the answer and the end marker are cut off, and it does not learn how to stop.

What you will do in the next lab

You turn the first conversation of the SFT data into the chat format and build the answer-only labels yourself. You count the ratio of tokens that go into the loss, and from the reference pretrained weights you train a masked model and an unmasked model under the same conditions. On held-out conversations, you compare the answer loss and question loss of the two models, and try extracting answers to held-out questions.