MiniMind — Train a Small Language Model Yourself, End to End
A MiniMind layer is made of six decisions
In one line
MiniMind-3 is a decoder with the same backbone as Qwen3. One layer (block) is Pre-Norm RMSNorm → attention (GQA, RoPE, QK-norm) → residual → RMSNorm → SwiGLU FFN → residual, and the embedding and output layer share weights. In this module, you shrink that configuration (768 dimensions and 8 layers to 128 dimensions and 4 layers) to build a model of about 1 million parameters, and check in numbers what each decision does.
Why this was needed
Quite a few of the decisions in the original Transformer (2017) have changed. Training was unstable (LayerNorm after the layer), inference used too much memory (K and V for every head), and it was hard to extend the length (added positional embeddings). Most public models today use the same remedies, and MiniMind shows those remedies in a few hundred lines of Python. You already computed attention by hand in the Transformer course, so here you look at how MiniMind's actual code implements those remedies, and at their numbers.
How it works
These are the defaults of MiniMindConfig and what they mean.
| Setting | MiniMind-3 | This course | Meaning |
|---|---|---|---|
| hidden_size | 768 | 128 | Length of the vector that represents one token |
| num_hidden_layers | 8 | 4 | Number of layers |
| num_attention_heads / num_key_value_heads | 8 / 4 | 4 / 2 | Q heads / K and V heads (GQA) |
| vocab_size | 6,400 | 1,024 | Vocabulary |
| intermediate_size | ceil(768·π/64)·64 = 2,432 | 448 | FFN middle width |
| rope_theta | 1e6 | 1e6 | Base of the RoPE rotation period |
RMSNorm. LayerNorm subtracts the mean and divides by the standard deviation. The RMSNorm paper argues that the re-centering of subtracting the mean is unnecessary and divides only by the root mean square (RMS). The MiniMind code likewise just multiplies x * rsqrt(mean(x²) + eps) by a weight. So the output's RMS returns to 1, but its mean does not become 0. Because it is Pre-Norm, normalizing before the layer, the residual path runs straight through without passing through normalization, and gradients flow well even as the network gets deeper.
GQA. There are 4 Q heads and 2 K and V heads. The output of k_proj is 2 × head_dim, half of Q's, and when computing, repeat_kv copies K and V twice each to match the number of Q heads. What stays in the cache at inference is the K and V before copying, so the KV cache per token is halved. The GQA paper reported that MQA, which reduces K and V heads to 1, loses quality, and that using a value in between gives quality close to MHA and speed close to MQA.
RoPE and QK-norm. Instead of adding position to the vector, it rotates q and k by their positions. If you rotate two vectors to their own positions and then take the dot product, only the difference of the two positions remains in the result — the dot products of (3, 7) and (103, 107) are the same. MiniMind normalizes q and k once more each with RMSNorm before rotating (q_norm and k_norm). This is the approach Qwen3 uses, and it keeps attention scores from growing too large.
SwiGLU. The FFN is down(silu(gate(x)) * up(x)). The gated GLU family has three matrices, so MiniMind sets the middle width to about π times hidden (rounded up to a multiple of 64). The GLU variants paper reported that such gated FFNs give better quality than ReLU and GELU FFNs.
Embedding sharing. Because tie_word_embeddings=True, lm_head.weight is embed_tokens.weight. It is a device that halves the vocabulary's share in a small model.
MoE variant. With use_moe=True, the FFN becomes 4 experts and only 1 is used per token. This is why MiniMind-3-MoE is "198M-A64M" — the parameters grow by more than three times, but the computation one token uses is almost the same as a dense model.
What it looks like in the field
You should be able to compute the parameter count, KV cache size, and serving memory from a single configuration file (config.json) on a model card. With "8 KV heads, head_dim 128, 32 layers," how many KB one token takes comes out right away, and that number decides the number of concurrent users. Conversely, if the parameter count you computed by formula differs from the model, you missed the shared weights or the normalization weights. Checking whether the loss before training is near ln(vocabulary) is also a common check — if initialization is broken, you can tell from the first step.
The MiniMind README cites the MobileLLM paper on how to split width and depth in a small model — for the same parameters, narrow and deep is generally better than shallow and wide, but below 512 dimensions the representation bottleneck becomes clear. The 128 dimensions of this course are well below that boundary. The size was chosen deliberately so that training finishes in a few minutes on two CPUs, and the price is measured in the last module.
What you will do in the next lab
You build a MiniMind model with the small configuration and count the parameters both by formula and in the model. You then check in turn GQA's projection size and KV cache per token, RMSNorm's output, RoPE's relative-position property, the loss before training, and the total and active parameters when you switch to MoE.