MiniMind — Train a Small Language Model Yourself, End to End
DPO teaches 'this one is better' without a reward model
In one line
This is the stage that teaches a model that has learned how to answer through SFT which answer is better. DPO (Direct Preference Optimization) does this with only pairs of a good answer (chosen) and a bad answer (rejected) to the same question, with a single classification loss, without a reward model and without reinforcement learning. MiniMind's train_dpo.py implements that loss in ten lines. In this module, you run DPO with pairs that have the same facts and differ only in speech style (입니다, the polite ending, versus 이다, the plain ending), and measure both whether the speech style changes and whether the content is preserved.
Why this was needed
RLHF first trains a reward model from human preferences, and then pushes the model with reinforcement learning (PPO) to maximize that reward without drifting too far from the original model. You have to run four copies of the model (policy, reference, reward, and value), and training is unstable. The DPO paper solves the optimum of this objective in closed form, so the same problem can be solved with two copies, policy and reference, and a simple classification loss. The MiniMind README says that DPO is an off-policy method that goes over static preference data several times, so it is stable, but because it does not explore on its own, it fits alignment such as preference and safety better than "the ability to get problems right." The experiment in this module shows exactly that boundary.
How it works
The log probability of one sentence is the sum of log p at each answer token. MiniMind keeps only the answer tokens with the same mask as SFT.
log_probs = torch.gather(F.log_softmax(logits, dim=2), 2, labels.unsqueeze(2)).squeeze(-1)
seq = (log_probs * mask).sum(dim=1) # 문장마다 답 토큰의 합
chosen, rejected = seq[:B // 2], seq[B // 2:] # 배치 앞 절반이 chosen, 뒤 절반이 rejected
logits = (π_chosen − π_rejected) − (ref_chosen − ref_rejected)
loss = −F.logsigmoid(β * logits).mean()
- Reference model. Load one more copy of the SFT model and freeze it (
eval()andrequires_grad_(False)). The more the policy comes to like chosen relatively more than the reference does, the lower the loss. If the policy equals the reference, the part in the parentheses is 0, so the loss is −log σ(0) = ln 2 ≈ 0.693. - β. It decides how far it may move from the reference. The paper interprets β·log(π/π_ref) as an implicit reward. MiniMind's default is 0.15.
- Learning rate. MiniMind's default is 4e-8, and a code comment says "5e-8 or below recommended, to avoid forgetting." It is hundreds of times smaller than even SFT (1e-5). Preference learning must push the model only a little.
One pitfall. The DPO loss looks only at the difference between chosen and rejected. So if the learning rate is large, it is common to pull rejected down a lot while pulling chosen down along with it — the difference widens and the loss falls, but the answer the model actually gives becomes neither one. With this course's model, if you raise lr to 1e-4, chosen's average log probability collapses from −3.5 to −16.5, and the rate of answering with the polite ending actually dropped. So do not look at the loss alone; look at the generated outputs and whether other abilities were preserved together.
What it looks like in the field
When feedback piles up that an internal chatbot mixes in casual speech or answers a forbidden topic, you build preference pairs and run DPO. A common experience then is "the tone got fixed but accuracy dropped" (commonly called the alignment tax). So the standard procedure is to measure both the preference metrics (preference accuracy and the rate of the desired tone) and the original ability metric (factual accuracy) before and after DPO. Remember also that preference data cannot teach facts — a fact the model does not know does not newly appear no matter how much chosen you show it. Building preference pairs is itself a cost. It takes time for people to read two answers and choose, and each chooser has different criteria, so noise gets mixed into the data. That is why it is easier to interpret results if you first settle "what is preferred" in one sentence (for example, polite speech, no unfounded assertions), and start with pairs made to differ in only that one criterion — this is why the pairs in this lab differ in only one place, the speech style.
How this course differs from the original MiniMind
MiniMind's preference data is real human preference pairs drawn from DPO-En-Zh-20k, and the two answers in a pair differ in content. The pairs in this course have the same content and differ only in one place, the speech style. That way you can measure "what preference learning moves and what it keeps" with a single variable. The learning rate also differs — MiniMind uses 4e-8 for a 64M model going over long data, while this course uses 1e-5 for a 1M model running only 100 steps. The numbers differ but the principle is the same: much smaller than the SFT learning rate, and measure the original ability as well while the metrics improve.
What you will do in the next lab
With the SFT reference model, you measure the log probabilities of one preference pair, implement the DPO loss exactly as the paper's formula, and check that it gives ln 2 when the policy equals the reference. You measure the preference accuracy and polite rate before DPO, then measure how the same metrics change after 100 steps of DPO and whether the accuracy ignoring speech style was preserved.