MiniMind — Train a Small Language Model Yourself, End to End
Shift the style with DPO and check the content was preserved
Goal
Implement MiniMind train_dpo.py's loss (−logσ(β·Δ), computed from the sum of answer-token log probabilities) yourself, and run 100 steps of DPO loading the SFT reference model as two copies, policy and reference. The preference pairs have the same facts and differ only in speech style (chosen …입니다., rejected …이다., the Korean polite and plain endings). You measure both whether the speech style changes and whether the accuracy ignoring speech style is preserved.
Why it matters
DPO pushes a model with only preference pairs, without a reward model and reinforcement learning. The implementation is ten lines, but because the loss looks only at the "difference," if you push wrongly the loss goes down nicely while the model breaks — for example, by pulling chosen and rejected down together and only widening the difference. With this course's model, if you raise the learning rate tenfold, this actually happens. So preference learning is not judged by the loss alone. The point of this lab is to measure together, before and after, the preference metrics (the share of held-out pairs where chosen is more plausible, and the share answering in the desired speech style) and the original ability metric (factual accuracy).
Steps
- In /root/mm/dpo/dpolib.py, create a function that turns preference pairs into a batch and the sum of answer-token log probabilities (
seq_logp), and write the two log probabilities of the first pair, computed with the reference model, to /root/mm/dpo/logp.json. - Implement
dpo_loss(정책 chosen, 정책 rejected, 기준 chosen, 기준 rejected, beta)indpolib.py(the Korean words are placeholders for policy and reference). Also write the loss when policy equals reference (loss_policy_equals_ref) to the step 1logp.json. - Write the reference model's metrics before DPO (preference accuracy, polite rate, and the two mean log probabilities) to /root/mm/dpo/baseline.json.
- Run 100 steps of DPO with β 0.1, lr 1e-5, and 8 pairs per step, and save the policy to /root/mm/dpo/dpo.pth.
- Measure the same metrics with the policy after DPO and write them to /root/mm/dpo/after.json.
- Measure the accuracy ignoring speech style on the first 40 questions of the SFT data before and after, and write it to /root/mm/dpo/drift.json.
- In /root/mm/dpo/report.md, write the three sections
## 손실,## 무엇이 바뀌었나, and## 무엇을 못 하나(the Korean headings mean "Loss", "What changed", and "What it cannot do"), and include the polite rate after DPO and chosen's mean log probability.
Notes
- For the answer-token mask, use as they are the labels of
mmkit.sft_encode(tok, 대화)(the slots that are not -100; the placeholder is the conversation). Shift the logits by one slot so thatlogits[:, :-1]andlabels[:, 1:]line up. - Metric definitions — preference accuracy: the share of the 100 pairs in
dpo_val.jsonlwhere log p(chosen) > log p(rejected). Polite rate: the share that ends with입니다.(the Korean polite ending) when answering the questions of the first 40 pairs greedily. - 100 steps take about 10 seconds on a node. Raise the learning rate to 1e-4, run once more, and compare what happens to the step 5 metrics (grading uses the 1e-5 result).
- Common mistakes: not freezing the reference model or using the same object as the policy, flipping the loss sign, and computing the log probability as a mean rather than a sum (MiniMind uses the sum).
- Sources: train_dpo.py · lm_dataset.py — DPODataset · DPO paper
A sentence's log probability is the sum over answer tokens
In /root/mm/dpo/dpolib.py, create batch(쌍들) (the input and labels with the first half chosen and the second half rejected; the placeholder is the pairs) and seq_logp(모델, 입력, 라벨) (the sum of answer-token log p for each sentence; the placeholders are the model, the input, and the labels), and with the reference model (/opt/mm/ref/sft.pth), write the two values of the first pair of dpo.jsonl to /root/mm/dpo/logp.json as chosen and rejected.
After log_softmax, pull out only the correct token's value with torch.gather, and remove slots whose label is -100 by multiplying by a mask of 0 (you must change -100 to 0 before gather to avoid an index error). The two answers differ only in the polite ending and the plain ending, so the probability difference at that part is the difference between the two values.
The DPO loss exactly as the formula
Create dpo_loss(policy_chosen, policy_rejected, ref_chosen, ref_rejected, beta) in dpolib.py — the four arguments are per-sentence log probability sums (1-dimensional tensors), and it returns the mean of −logσ(β·[(π_c − ref_c) − (π_r − ref_r)]). With the first 8 pairs from the reference model, write the value when policy equals reference to loss_policy_equals_ref in logp.json.
If the policy equals the reference, the part in the parentheses is 0, so −log σ(0) = ln 2. The grader loads your dpo_loss and checks whether it matches the paper's formula on six random pairs and two values of β, and whether the loss is smaller than ln 2 when the policy likes chosen more (the sign). Using F.logsigmoid keeps it numerically stable.
Metrics before DPO
Write the script /root/mm/dpo/metrics.py (arguments: weights path, result path), and with the reference model measure the preference accuracy on the 100 pairs of dpo_val.jsonl (pref_acc), the polite rate on the questions of the first 40 pairs (polite_rate), and the mean log probabilities of chosen and rejected (mean_logp_chosen and mean_logp_rejected), and write them to /root/mm/dpo/baseline.json.
The SFT data was half and half in speech style, so the reference model mixes the two styles about equally. If you make the metric script take the weights path and the result path as arguments, you can use it as it is in step 5.
100 steps of DPO
Write the script /root/mm/dpo/train_dpo.py, load the SFT model as two copies, policy and reference, freeze the reference, draw 8 pairs at a time from dpo.jsonl, run DPO with β 0.1, lr 1e-5 (AdamW), and 100 steps (seed 42), and then save the policy's state_dict to /root/mm/dpo/dpo.pth.
Compute the reference model's log probabilities inside torch.no_grad(). MiniMind's default learning rate is 4e-8, but for this small model and short training, 1e-5 moves the speech style while preserving the content. The loss starts at ln 2 and goes down.
Metrics after DPO
Measure the same metrics as in step 3 with dpo.pth and write them to /root/mm/dpo/after.json. The grader checks whether the preference accuracy is 0.8 or higher and the polite rate rose by 20 percentage points or more compared to before DPO.
Also compare the two mean log probabilities before and after. Both chosen and rejected going down is common in DPO — the problem is when chosen goes down too much and the model gives an answer that is neither one.
Was the content preserved
Have the model answer greedily the questions of the first 40 single-turn conversations in sft.jsonl, strip the trailing 입니다. and 이다. (the Korean polite and plain endings), measure the share equal to the data's answer for the reference model (fact_acc_ref) and the DPO policy (fact_acc_dpo), and write it to /root/mm/dpo/drift.json.
If you strip the speech style and compare, only "what was said" remains. The preference pairs have the same facts and differ only in speech style, so in a well-done DPO this value should stay nearly the same. If it drops by more than 10 percentage points, it went too far from the reference.
The effect and limits of preference learning
In /root/mm/dpo/report.md, write the three sections ## 손실, ## 무엇이 바뀌었나, and ## 무엇을 못 하나 (the Korean headings mean "Loss", "What changed", and "What it cannot do"), and include the polite_rate and mean_logp_chosen from step 5 as numbers.
In the last section, write what cannot be taught with preference pairs (facts the model does not know) and what you saw when you raised the learning rate.