TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

The loss curve shows what a model learns first

Continue in TT Lab

In one line

Pretraining is repeating "guess the next token" millions of times. MiniMind's train_pretrain.py runs this loop with standard parts: AdamW, a cosine learning rate, and gradient clipping. In this module, you shrink that loop to the size of one CPU and write it yourself, and read the curve in which the loss falls from 6.9 to around 1 — where it drops sharply, where it slows, and how it breaks when the learning rate is large.

Why this was needed

A training loop mostly looks like the two lines loss.backward(); optimizer.step(). But anyone who has trained a model all the way even once knows that the decisions around those two lines decide the result. If the learning rate is large, the loss goes down and then jumps up; if you do not fix the seed, you cannot reproduce yesterday's result; and if you do not measure the validation loss, you cannot tell memorizing from learning. The loss curve is the only window in which all of this shows up. The MiniMind README says that pretraining the 64M model takes a little over an hour on a single 3090, and to make good use of that hour you must be able to read the curve of the first few minutes.

How it works

Boiled down to the essentials, MiniMind's loop looks like this.

for step, (input_ids, labels) in enumerate(loader, start=1):
    lr = get_lr(epoch * iters + step, epochs * iters, learning_rate)
    for g in optimizer.param_groups: g["lr"] = lr
    res = model(input_ids, labels=labels)
    loss = (res.loss + res.aux_loss) / accumulation_steps   # aux_loss 는 MoE 균형 손실, 밀집 모델은 0
    scaler.scale(loss).backward()
    if step % accumulation_steps == 0:
        scaler.unscale_(optimizer)
        torch.nn.utils.clip_grad_norm_(model.parameters(), grad_clip)   # 기본 1.0
        scaler.step(optimizer); scaler.update(); optimizer.zero_grad(set_to_none=True)

Learning rate. get_lr is lr × (0.1 + 0.45 × (1 + cos(π·t/T))). It starts at lr as is and, at the end, comes down to 0.1·lr along a cosine. It is characteristic that it does not go all the way down to 0. The default lr is 5e-4, but the small model in this course sets a much larger 3e-3 — the smaller the model, the larger a learning rate it can tolerate.

Gradient accumulation and clipping. With the default batch of 32 and accumulation of 8, you actually see 256 samples and take one step. clip_grad_norm_ shrinks the gradient vector to 1 if its overall length exceeds 1 — so that one occasional spiking batch does not wreck the weights. On a GPU it uses bfloat16 autocast, but on a CPU it is nullcontext(), so it runs in fp32.

Reading the curve. Our model starts at ln 1024 ≈ 6.93 and drops below 3 within a few dozen steps. In this steep stretch, the model learns which tokens appear often. The loss of a model that looks at no context and knows only frequency is the unigram entropy (about 4.6 nats on our corpus). The curve going below that means it has begun to use the preceding token to narrow down the next one. After that it is slow — long rules such as sentence patterns, particles, and pairs of a village and its local specialty are learned slowly.

Reproducibility. You must fix both the seed for weight initialization and the seed for drawing batches. MiniMind's setup_seed fixes random, numpy, and torch all at once. With the same CPU and the same number of threads, the loss over 20 steps comes out without a single digit of difference.

What it looks like in the field

Before running training for hours, there are things to check with a short training of a few minutes. Is the first loss near ln(vocabulary)? Does it drop sharply within a few dozen steps? Does the validation loss follow the training loss? If the first loss is strange, the data or labels are wrong, and if it does not drop at all, the learning rate is too small or the gradients are not flowing. Conversely, if the loss drops and then spikes or sticks at some value and does not move, the learning rate is too large — in this lab you raise lr to 0.05 and see it yourself.

A low loss does not make a good model either. Our corpus consists of sentences stamped out from a template, so the next token is almost determined and the loss falls to around 1. With a real corpus, a model of the same size would not fall this far. Numbers can be compared only within the same data and the same tokenizer.

What you will do in the next lab

You write train.py modeled on MiniMind's loop and run it twice with the same seed to see whether it differs by even one digit. You train 300 steps and leave a checkpoint and a loss log, extract the shape of the curve as numbers, compare it with the unigram entropy, see what a learning rate of 0.05 breaks, and use the trained model to continue a text.