TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

Run SFT in chat format and measure what loss masking changes

Continue in TT Lab

Goal

Cut conversations in MiniMind's chat format (<|im_start|>역할\n…<|im_end|>\n, where the Korean word stands for the role) and build answer-only (assistant) labels yourself, like generate_labels. From the reference pretrained weights, train a masked model and an unmasked model under the same conditions, and see in numbers how the loss on answer tokens and on question tokens diverges on held-out conversations.

Why it matters

SFT is not a new algorithm but the same next-token prediction. The only difference is what you treat as the correct answer. If you treat the whole conversation as the answer, the model splits its learning signal with imitating the question, and after the answer ends it may make up the next question itself. If you treat only the answer as correct, the question becomes a condition, and you must put the end marker in the correct answer too for it to learn when to stop. This boundary is decided by a few -100 entries in the label array, and no error is raised even if it is wrong. So printing the labels yourself and measuring with two models what masking actually changes is the most reliable check.

Steps

  1. Convert the first conversation of /opt/mm/data/sft.jsonl into a chat-format string and save it to /root/mm/sft/sample0.txt.
  2. Create encode(대화, max_len=64, mask=True) in /root/mm/sft/sftlib.py (the placeholder is the conversation), and save the input_ids and labels of the first 20 conversations to /root/mm/sft/labels.json.
  3. Write the real token count excluding padding, the count of tokens with labels remaining, and their ratio over the whole SFT data to /root/mm/sft/ratio.json.
  4. From the reference pretrained weights (/opt/mm/ref/pretrain.pth), do SFT for 300 steps with the masked labels and save /root/mm/sft/sft_masked.pth.
  5. Under the same conditions, put the entire input in the labels (only padding is -100) and save /root/mm/sft/sft_nomask.pth.
  6. On the held-out conversations (sft_val.jsonl), write the answer-token loss and the question-side token loss of the two models to /root/mm/sft/compare.json.
  7. With the masked model, extract greedy answers to 20 held-out questions and save them to /root/mm/sft/answers.jsonl.
  8. In /root/mm/sft/report.md, write the three sections ## 채팅 형식, ## 손실 마스킹, and ## 한계 (the Korean headings mean "Chat format", "Loss masking", and "Limitations"), and include the masked model's question loss and the ratio from step 3.

Notes

A conversation as one line

Convert the conversations of the first line of /opt/mm/data/sft.jsonl into the MiniMind chat format (<|im_start|>역할\n내용<|im_end|>\n for each turn, where the Korean words stand for the role and the content) and save it to /root/mm/sft/sample0.txt.

mmkit.chat_text(대화) (the placeholder is the conversation) makes this format. You may also make it yourself — a newline after the role, and <|im_end|> and a newline after the content. Do not include the <think> tag for thinking mode or a system prompt.

Labels that keep only the answer

Create encode(대화, max_len=64, mask=True) in /root/mm/sft/sftlib.py (the placeholder is the conversation) — it returns input_ids, the chat format cut with /opt/mm/ref/tokenizer.json, truncated to 64 tokens and filled with padding (0), and labels, which are the original tokens only from after the answer's start marker to the end marker and -100 everywhere else. Save the results for the first 20 conversations to /root/mm/sft/labels.json as [{"input_ids": […], "labels": […]}, …].

Like MiniMind's generate_labels, scan the token sequence to find the stretch equal to the start marker pieces, and fill the labels from there until the end marker pieces end. A two-turn conversation has two such stretches. The grader compares digit by digit with labels made by the same rule.

What percent of tokens go into the loss

Convert all of sft.jsonl with encode, and write the number of non-padding input tokens (real_tokens), the number of labels that are not -100 (label_tokens), and their ratio (ratio) to /root/mm/sft/ratio.json.

If the question and role markers are only conditions, the learning signal comes only from the rest. Even with the same number of steps as pretraining, the tokens SFT actually learns are this ratio.

Masked SFT

Write the script /root/mm/sft/train_sft.py (arguments --nomask and --out), start from /opt/mm/ref/pretrain.pth, do SFT with the masked labels for 300 steps (batch 16, length 64, lr 1e-3, MiniMind cosine, seed 42), and save it to /root/mm/sft/sft_masked.pth.

Only the data changes from the loop in the earlier module — model(X[ix], labels=Y[ix]). The grader checks whether the answer loss on held-out conversations is below 1.2 and whether the question-side loss remains high (3 or more). If you masked, it is because the question was not learned.

SFT without masking

Keep all conditions the same as in step 4 and change only the labels to input_ids as they are (only padding slots are -100), train, and save to /root/mm/sft/sft_nomask.pth.

If you make one argument produce both, as in encode(대화, mask=False) (the placeholder is the conversation), the code guarantees the conditions are the same. The grader checks whether this model's question-side loss is low (below 1.5) — it is evidence that it learned the question too.

Measure the two models on held-out conversations

Feed all the conversations of sft_val.jsonl into the two models, and write the loss keeping only the answer-stretch labels (answer_loss) and the loss keeping only the non-answer slots (excluding padding) (prompt_loss) to /root/mm/sft/compare.json as {"masked": {…}, "nomask": {…}}.

For the same input, mask the labels two ways and compute cross_entropy(…, ignore_index=-100) twice. Compare logits shifted by one slot (logits[:, :-1] and labels[:, 1:]). Also compare the answer loss of the two models at the same number of steps.

Answer the held-out questions

With the masked model, extract greedy answers to the first question of the first 20 conversations of sft_val.jsonl and save them to /root/mm/sft/answers.jsonl, one per line, as {"q": 질문, "a": 답, "ref": 자료의 답} (the placeholders are the question, the answer, and the answer in the data).

mmkit.greedy(model, tok, 질문) (the placeholder is the question) generates by attaching <|im_start|>assistant\n to the chat format and stops at <|im_end|>. The held-out questions are in question shapes never seen in SFT, so it is normal for answers to come out with the right format but wrong content — you measure why in the last module.

Write down the effect of masking

In /root/mm/sft/report.md, write the three sections ## 채팅 형식, ## 손실 마스킹, and ## 한계 (the Korean headings mean "Chat format", "Loss masking", and "Limitations"), and include the masked model's prompt_loss from step 6 and the ratio from step 3 as numbers.

Write one line on why a high question loss is a "good" signal. In the limitations section, write the shape of the wrong answers from step 7.