MiniMind — Train a Small Language Model Yourself, End to End
Change only the style with MiniMind's LoRA and confirm the base weights are unchanged
Goal
Attach a rank-8 LoRA to the SFT reference model with apply_lora, save_lora, and load_lora of MiniMind's model_lora.py. Freeze the base weights and train only LoRA to change the answers to a "…da-nyang." speech style, and check that the base weights did not change by even one bit and that the merged model gives the same output.
Why it matters
LoRA makes two promises — train only a few parameters, and leave the base model as it is. The former reduces memory and storage, and the latter lets you swap several LoRAs on one base. But if you forget to freeze, both promises break silently. Training goes well and the speech style changes, but the base weights move too, so all other LoRAs go wrong. That is why this lab does not look only at the result (the speech style), but compares the base weights after training with the original one tensor at a time. You also confirm in code that MiniMind's implementation has no alpha scaling and attaches only to Linear layers with equal dimensions.
Steps
- Load
/opt/mm/ref/sft.pth, runapply_lora(model, rank=8), and write the names of the modules that LoRA attached to into /root/mm/lora/targets.json asrankandtargets. - Write the maximum logit difference between before and right after attaching to /root/mm/lora/zero.json as
max_abs_diff. - Freeze every parameter that has no
lorain its name, and write the number of trainable parameters, the total number, and the ratio to /root/mm/lora/params.json. - Train only LoRA for 150 steps with
lora_nyang.jsonland save /root/mm/lora/lora.pth withsave_lora. - From the trained model, save the weights with LoRA removed to /root/mm/lora/base_after.pth. They must equal the original SFT weights.
- On the held-out questions (
lora_nyang_val.jsonl), measure the fraction of answers ending in "nyang." without LoRA and with LoRA, and write it to /root/mm/lora/effect.json. - Save the weights merged as W + B·A to /root/mm/lora/merged.pth.
- In /root/mm/lora/report.md, write the three sections
## 어디에 붙나,## 무엇이 바뀌나, and## 합치기(the Korean headings mean "Where it attaches", "What changes", and "Merging"), and include the number of trained parameters and the "nyang" rate after attaching LoRA.
Notes
from model.model_lora import apply_lora, save_lora, load_lora— open/opt/minimind/model/model_lora.pyyourself. It is 60 lines.- For SFT labels,
mmkit.sft_encode(tok, 대화)(the placeholder is the conversation) makes them by the same rule as the one built in the SFT module. - 150 steps take about 10 seconds on a node. Generating the 60 held-out questions takes about 5 seconds for two runs.
- Common mistakes: passing all of
model.parameters()to the optimizer without freezing withrequires_grad=False(the base weights move too), using memory to compute gradients of base parameters even though they are frozen, and saving the full state_dict instead of usingsave_lora. - Sources: model_lora.py · train_lora.py · LoRA paper
Where LoRA attaches
Load with mmkit.load_model("/opt/mm/ref/sft.pth"), run apply_lora(model, rank=8), and write the names of the modules that gained a lora attribute, in order, to /root/mm/lora/targets.json as {"rank": 8, "targets": [...]}.
Loop over model.named_modules() and collect those with hasattr(m, 'lora'). apply_lora attaches only to a Linear with in_features == out_features — first guess which projections in our model satisfy that condition.
Right after attaching, nothing changes
Compute the logits for the first 64 tokens of /opt/mm/ref/val.npy before and right after apply_lora, and write the maximum absolute difference to /root/mm/lora/zero.json as max_abs_diff.
MiniMind's LoRA initializes A with a normal distribution and B with 0. Since B·A is 0, the side branch's output is 0, and training learns "how far to move away from the base model" starting from 0.
Freeze so that only LoRA trains
Start writing the training script /root/mm/lora/train_lora.py — in the model with LoRA attached, set parameters without lora in their name to requires_grad=False and those with it to True, and write the number of trainable parameters (trainable), the total number (total), and the ratio (ratio) to /root/mm/lora/params.json.
It is exactly the rule MiniMind's train_lora.py uses. Each place has A (8×128) + B (128×8), and you multiply by the number of places attached. The total includes the LoRA parameters too.
Learn the speech style with LoRA
Convert /opt/mm/data/lora_nyang.jsonl with mmkit.sft_encode, pass only the LoRA parameters to AdamW (lr 5e-3) and train 150 steps (batch 16, seed 42), and save /root/mm/lora/lora.pth with save_lora(model, "/root/mm/lora/lora.pth").
save_lora saves only …lora.A.weight and …lora.B.weight of each LoRA-attached module as fp16. Compare the file size with the base model (a little over 4MB). The grader checks whether the keys are only LoRA and whether B moved from 0.
Are the base weights unchanged
From the trained model's state_dict, remove those with .lora. in their name and save the rest to /root/mm/lora/base_after.pth. The grader compares it tensor by tensor with the original /opt/mm/ref/sft.pth.
If freezing was done properly, it does not differ by even one bit. If you did not turn off requires_grad and passed all parameters to the optimizer, the base weights also receive gradients and move, and you get caught here. The PyTorch optimizer skips parameters with no gradient (None), so as long as freezing is done properly, the base stays as it is whatever you passed to the optimizer.
How much did the speech style change
Have the model answer the 60 questions in lora_nyang_val.jsonl greedily, measure the fraction of answers ending in "nyang." for the base model (base_rate) and for the model with LoRA attached (lora_rate), and write it to /root/mm/lora/effect.json.
Make the model with LoRA attached by loading the base model, running apply_lora(rank=8), and then load_lora(model, 경로) (the placeholder is the path). The held-out questions are combinations never seen during LoRA training, so you check whether the speech style was learned as a "way of answering" rather than a "question shape".
Merge to remove the side branch
Using the A and B in lora.pth, add B @ A to the weights of the corresponding modules (without an alpha scaling) and save the merged state_dict to /root/mm/lora/merged.pth. The grader puts these weights into a model without LoRA and checks whether it gives the same logits as the model with LoRA attached.
A has shape (rank, input) and B has (output, rank), so B @ A is (output, input) — the same shape as the original weight. MiniMind's merge_lora does the same and saves in fp16. Here, we save in fp32 for comparison.
A record of confirming LoRA's promises
In /root/mm/lora/report.md, write the three sections ## 어디에 붙나, ## 무엇이 바뀌나, and ## 합치기 (the Korean headings mean "Where it attaches", "What changes", and "Merging"), and include the trainable from step 3 and the lora_rate from step 6 as numbers.
Also try writing one line on what fraction of the base model the LoRA file is, and, since the implementation has no alpha scaling, what you must look at again when you change the rank.