TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

DPO teaches 'this one is better' without a reward model

Continue in TT Lab

In one line

This is the stage that teaches a model that has learned how to answer through SFT which answer is better. DPO (Direct Preference Optimization) does this with only pairs of a good answer (chosen) and a bad answer (rejected) to the same question, with a single classification loss, without a reward model and without reinforcement learning. MiniMind's train_dpo.py implements that loss in ten lines. In this module, you run DPO with pairs that have the same facts and differ only in speech style (입니다, the polite ending, versus 이다, the plain ending), and measure both whether the speech style changes and whether the content is preserved.

Why this was needed

RLHF first trains a reward model from human preferences, and then pushes the model with reinforcement learning (PPO) to maximize that reward without drifting too far from the original model. You have to run four copies of the model (policy, reference, reward, and value), and training is unstable. The DPO paper solves the optimum of this objective in closed form, so the same problem can be solved with two copies, policy and reference, and a simple classification loss. The MiniMind README says that DPO is an off-policy method that goes over static preference data several times, so it is stable, but because it does not explore on its own, it fits alignment such as preference and safety better than "the ability to get problems right." The experiment in this module shows exactly that boundary.

How it works

The log probability of one sentence is the sum of log p at each answer token. MiniMind keeps only the answer tokens with the same mask as SFT.

log_probs = torch.gather(F.log_softmax(logits, dim=2), 2, labels.unsqueeze(2)).squeeze(-1)
seq = (log_probs * mask).sum(dim=1)            # 문장마다 답 토큰의 합
chosen, rejected = seq[:B // 2], seq[B // 2:]  # 배치 앞 절반이 chosen, 뒤 절반이 rejected
logits = (π_chosen − π_rejected) − (ref_chosen − ref_rejected)
loss = −F.logsigmoid(β * logits).mean()

One pitfall. The DPO loss looks only at the difference between chosen and rejected. So if the learning rate is large, it is common to pull rejected down a lot while pulling chosen down along with it — the difference widens and the loss falls, but the answer the model actually gives becomes neither one. With this course's model, if you raise lr to 1e-4, chosen's average log probability collapses from −3.5 to −16.5, and the rate of answering with the polite ending actually dropped. So do not look at the loss alone; look at the generated outputs and whether other abilities were preserved together.

What it looks like in the field

When feedback piles up that an internal chatbot mixes in casual speech or answers a forbidden topic, you build preference pairs and run DPO. A common experience then is "the tone got fixed but accuracy dropped" (commonly called the alignment tax). So the standard procedure is to measure both the preference metrics (preference accuracy and the rate of the desired tone) and the original ability metric (factual accuracy) before and after DPO. Remember also that preference data cannot teach facts — a fact the model does not know does not newly appear no matter how much chosen you show it. Building preference pairs is itself a cost. It takes time for people to read two answers and choose, and each chooser has different criteria, so noise gets mixed into the data. That is why it is easier to interpret results if you first settle "what is preferred" in one sentence (for example, polite speech, no unfounded assertions), and start with pairs made to differ in only that one criterion — this is why the pairs in this lab differ in only one place, the speech style.

How this course differs from the original MiniMind

MiniMind's preference data is real human preference pairs drawn from DPO-En-Zh-20k, and the two answers in a pair differ in content. The pairs in this course have the same content and differ only in one place, the speech style. That way you can measure "what preference learning moves and what it keeps" with a single variable. The learning rate also differs — MiniMind uses 4e-8 for a 64M model going over long data, while this course uses 1e-5 for a 1M model running only 100 steps. The numbers differ but the principle is the same: much smaller than the SFT learning rate, and measure the original ability as well while the metrics improve.

What you will do in the next lab

With the SFT reference model, you measure the log probabilities of one preference pair, implement the DPO loss exactly as the paper's formula, and check that it gives ln 2 when the policy equals the reference. You measure the preference accuracy and polite rate before DPO, then measure how the same metrics change after 100 steps of DPO and whether the accuracy ignoring speech style was preserved.