TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

Shrink the MiniMind config into a model and check each layer decision with numbers

Continue in TT Lab

Goal

Shrink MiniMindConfig to 128 dimensions, 4 layers, 4 Q heads, 2 KV heads, and a vocabulary of 1024 to build a model, and match the parameter count between formula and model. Measure directly with the MiniMind code GQA's KV cache savings, RMSNorm's output, RoPE's relative-position property, the loss before training, and MoE's total and active parameters.

Why it matters

A single configuration file decides the model's size and serving cost. If you can count parameters by formula, you know right away whether you missed shared weights or normalization, and if you can compute the KV cache in bytes per token, you can compute how many users one card can take. And the remedies that today's models use (RMSNorm, GQA, RoPE, SwiGLU, embedding sharing) each have a reason. Instead of trusting the explanations, this lab confirms those properties in numbers with MiniMind's actual code — normalization that does not subtract the mean, K and V cut in half, and a rotation in which only distance remains.

Steps

  1. Save the small configuration to /root/mm/arch/config.json (hidden_size 128, num_hidden_layers 4, vocab_size 1024, num_attention_heads 4, num_key_value_heads 2, max_position_embeddings 512).
  2. Compute the parameter count by formula and also count it in the model, and write it to /root/mm/arch/count.json as intermediate_size, per_layer, formula, and model.
  3. Write to /root/mm/arch/gqa.json the output sizes of the first layer's attention q_proj and k_proj, and the bytes to cache one token in fp32 (as GQA is / if the KV heads were equal to the number of Q heads).
  4. After torch.manual_seed(0), feed randn(4, 128)*5+3 into the first layer's input_layernorm, and write the input and output RMS and the output mean to /root/mm/arch/rmsnorm.json.
  5. After torch.manual_seed(0), draw one q and one k, place them at positions (3,7), (103,107), and (3,50) with MiniMind's apply_rotary_pos_emb, and write the dot products to /root/mm/arch/rope.json.
  6. Measure the loss of the pre-training model made with mmkit.seed_all(0) on the first 16×128 tokens of the validation data, and write it to /root/mm/arch/init_loss.json as loss and ln_vocab.
  7. Write the total and active parameters of a model with use_moe=True added to the same configuration to /root/mm/arch/moe.json.
  8. In /root/mm/arch/report.md, write the three sections ## 파라미터는 어디에, ## GQA 와 KV 캐시, and ## RoPE 와 RMSNorm (the Korean headings mean "Where the parameters are", "GQA and the KV cache", and "RoPE and RMSNorm"), and include the parameter count from step 2 and the GQA cache bytes per token from step 3.

Notes

One small configuration file

Write hidden_size 128, num_hidden_layers 4, vocab_size 1024, num_attention_heads 4, num_key_value_heads 2, and max_position_embeddings 512 to /root/mm/arch/config.json. Use the MiniMindConfig defaults for the remaining values.

MiniMind-3 is 768, 8, 6400, 8, and 4. If you reduce the layers and dimension, the parameters shrink with the square of the dimension, so it trains in a few minutes even on a CPU. The grader actually builds a model from this file.

Count parameters by formula

Compute the parameter count by formula (formula), also count it in the model (model), and write it to /root/mm/arch/count.json as intermediate_size, per_layer, formula, and model. The two values must be equal.

One layer = the q, k, v, and o projections + q_norm and k_norm (head_dim each) + gate, up, and down (hidden×intermediate each) + two RMSNorms (hidden each). To this, add the embedding (vocabulary×hidden, counted once because it is shared with the output layer) and the final RMSNorm. intermediate_size is ceil(hidden·π/64)·64.

What GQA reduces

Write to /root/mm/arch/gqa.json the first layer's attention q_proj.out_features, k_proj.out_features, and n_rep, and the bytes to cache one token in fp32 (both K and V × layers × KV heads × head_dim × 4) — as GQA is (kv_bytes_per_token_gqa) and when the KV heads equal the number of Q heads (kv_bytes_per_token_mha).

MiniMind's Attention creates only num_key_value_heads K and V heads and, when computing, copies them n_rep times with repeat_kv. The cache (past_kv) holds the K and V from before the copy.

RMSNorm does not subtract the mean

After torch.manual_seed(0), create x = torch.randn(4, 128) * 5 + 3 and feed it into the first layer's input_layernorm, and write the RMS of the input and output (the mean over rows of sqrt(mean(x²))) and the mean of the whole output to /root/mm/arch/rmsnorm.json as rms_before, rms_after, and mean_after.

MiniMind's RMSNorm multiplies x * rsqrt(mean(x²) + eps) by a weight (initially 1). The RMS returns to 1, but since the mean is not subtracted, the +3 of the input remains as a trace in the output — with LayerNorm, the mean would be 0.

RoPE leaves only distance

After torch.manual_seed(0), draw q = torch.randn(1,1,1,head_dim) and k = torch.randn(1,1,1,head_dim), rotate q to the first position and k to the second position using the model's freqs_cos, freqs_sin, and apply_rotary_pos_emb, and write the dot products for (3,7), (103,107), and (3,50) to /root/mm/arch/rope.json as dot_3_7, dot_103_107, and dot_3_50.

apply_rotary_pos_emb(q, k, cos[p:p+1], sin[p:p+1]) rotates q and k to the same position, so call q and k separately and rotate each to its own position. The dot products of two pairs with the same distance of 4 must match to the fifth decimal place.

The loss before training is ln(vocabulary)

After fixing the seed with mmkit.seed_all(0), create a new model and write the loss (model(x, labels=x).loss) on the first 16×128 tokens of /opt/mm/ref/val.npy (view(16, 128)) to /root/mm/arch/init_loss.json as loss and ln_vocab.

A randomly initialized model gives almost the same probability (1/1024) to every token. Its cross entropy is ln 1024 ≈ 6.93. If you are far from this value, the initialization or the label handling is wrong.

MoE is big, but one token uses little

Write to /root/mm/arch/moe.json the total parameters of the model with use_moe=True added to the step 1 configuration (default 4 experts, 1 per token), and the parameters one token actually uses (total − size of one expert × number of experts + size of one expert × experts per token), as num_experts, top_k, total_params, and active_params.

The size of one expert is the sum of the parameters whose names contain mlp.experts.0. (one per layer, so the sum of expert 0 across all layers). MiniMind's trainer_utils.get_model_params prints "198M-A64M" with the same calculation.

A report explaining the configuration

In /root/mm/arch/report.md, write the three sections ## 파라미터는 어디에, ## GQA 와 KV 캐시, and ## RoPE 와 RMSNorm (the Korean headings mean "Where the parameters are", "GQA and the KV cache", and "RoPE and RMSNorm"), and include the parameter count from step 2 (model) and kv_bytes_per_token_gqa from step 3 as numbers.

In the first section, write what percent each of the FFN, attention, and embedding is; in the second, how many times GQA reduced the cache; and in the third, one line each on the properties you saw in steps 4 and 5.