MiniMind — Train a Small Language Model Yourself, End to End
LoRA adds a side branch without touching the weights
In one line
LoRA freezes the trained weight W, attaches two small matrices next to it (A: input→r, B: r→output), and trains only B·A. B starts at 0, so at first the model does not change at all, and when training ends you can merge it into W + B·A and use it without the side branch. MiniMind's model_lora.py implements this in 60 lines. In this module, you attach LoRA to an SFT model to change its speech style and directly check that the base weights did not change by even one bit.
Why this was needed
The LoRA paper uses GPT-3 175B as an example — if you keep a separate fully fine-tuned model for each task, one task needs 175 billion parameters. It reported that if you freeze the pretrained weights and train only low-rank matrices, the trainable parameters drop to one ten-thousandth, GPU memory drops to a third, and unlike adapters, it adds no inference latency. The MiniMind README cites the same use — leave the base model's general ability alone and attach a domain such as medicine, or a self-identity such as "who am I," with LoRA. If you have enough data, full SFT also works, but then you separately need to mix in data so that you do not overfit to the domain data and lose general ability.
How it works
MiniMind's implementation has three parts.
class LoRA(nn.Module):
def __init__(self, in_features, out_features, rank):
self.A = nn.Linear(in_features, rank, bias=False) # 정규분포(std 0.02)
self.B = nn.Linear(rank, out_features, bias=False) # 0
def forward(self, x): return self.B(self.A(x))
def apply_lora(model, rank=16):
for name, module in model.named_modules():
if isinstance(module, nn.Linear) and module.in_features == module.out_features:
module.lora = LoRA(...); module.forward = 원래 forward(x) + lora(x)
Where it attaches. It attaches only to a Linear whose input and output dimensions are the same. In our model (128 dimensions, Q heads 4×32), q_proj (128→128) and o_proj (128→128) qualify, while k_proj and v_proj (128→64) and the FFN (128↔448) are left out. 4 layers × 2 = 8 places, and each place has A (8×128) + B (128×8) = 2,048, so the trainable parameters are 16,384, 1.6% of the total. The paper treats the question of which matrices to attach to separately, and MiniMind settled it with this simple rule.
Starting point. Since B is 0, B·A = 0. The output right after attaching is exactly the same as the base model. Training learns "how far to move away from the base model" starting from 0.
Scaling. The paper multiplies ΔW·x by α/r so that you do not have to pick the learning rate again when you change the rank. MiniMind's implementation adds without this scaling — so when you change the rank, you must look at the learning rate again too.
Freezing and saving. train_lora.py sets every parameter without lora in its name to requires_grad=False, and passes only the LoRA parameters to the optimizer. For saving, save_lora extracts only the LoRA weights as fp16 — tens of KB. To use it, you do apply_lora on the base model and then load_lora, or build a single model merged with merge_lora (W + B·A).
What it looks like in the field
If you keep a different LoRA for each customer or task on one base model, storage is one base plus several small files, and serving just swaps the LoRA per request. If the base weights are secretly changed in this setup, all the LoRAs go wrong at once — a single training script that forgot to freeze can cause that. So when training ends, comparing the base weights against the original bit for bit is a cheap and reliable check. Conversely, if you will use it for only one purpose, merge it before deployment to remove the side-branch computation at inference.
How this course differs from the original MiniMind
MiniMind's default rank is 16, and train_lora.py runs 10 epochs at lr 1e-4. This course uses rank 8, lr 5e-3, and 150 steps. The model is small and the task (a single speech style) is simple, so reducing the rank is enough, and we raised the learning rate to see the effect in a short time. Since it is an implementation without an alpha scaling, the rank and the learning rate move together — if you double the rank, the same learning rate effectively becomes a larger step, as much as the size of B·A grows. Also, MiniMind takes attaching LoRA on top of the SFT model (full_sft) as the default. If you attach directly to a pretrained model, it has to learn the conversation format first, and 16,384 parameters are not enough for that.
What you will do in the next lab
You attach a rank-8 LoRA to the SFT reference model and check where it attaches and whether the output right after attaching stays the same. After freezing so that only LoRA trains, you train 150 steps on data in a "…da-nyang." speech style, and measure whether the base weights are unchanged, how much the speech style changed on held-out questions, and whether the merged model gives the same output as the model with LoRA attached.