MiniMind — Train a Small Language Model Yourself, End to End
Padding wastes compute; packing blurs document boundaries
In one line
Pretraining data is made into "one long line of tokens" and cut into batches to feed the model. MiniMind keeps one document per row and pads it to the maximum length — it is simple and documents do not get mixed, but if documents are short, most of the computation is thrown away on padding. Packing, which concatenates several documents to fill a row completely, throws away no computation, but within one window the earlier document is visible to the later one. This module measures the cost of both approaches in numbers on our corpus.
Why this was needed
The model receives an integer tensor of shape (배치, 길이), that is, batch by length. Document lengths vary, but a tensor must be rectangular, so you have to even them out somewhere. There are two ways.
Padding. Put one document on one row and fill the missing slots with padding tokens. MiniMind's PretrainDataset does it this way.
tokens = tokenizer(text, max_length=self.max_length - 2, truncation=True).input_ids
tokens = [bos_token_id] + tokens + [eos_token_id]
input_ids = tokens + [pad_token_id] * (self.max_length - len(tokens))
labels = input_ids.clone(); labels[input_ids == pad_token_id] = -100
Padding slots have a label of -100 so they do not enter the loss, but the computation is done all the same. A matrix multiplication does not know what is padding. The documents in our corpus average 33 tokens, and if you set the maximum length to 128, over 70% of the computation goes to padding. Conversely, if you set the maximum length short, long documents get cut. The MiniMind README also records this tradeoff — short samples waste computation on padding, and long samples are cut and lose information. So it lists a recommended max_seq_len separately for each dataset.
Packing. Wrap each document in [bos] … [eos], join them into one long line, and cut that line to a fixed length. There are no wasted slots. In exchange, a window holds three or four documents, and since the causal mask lets each token see "all the preceding tokens," tokens of a later document see the earlier document. Most pretraining accepts this — [eos] and [bos] mark the boundary, and the model learns to ignore what is beyond the boundary. To stop it from crossing the boundary you have to build a per-document attention mask separately, and that makes the implementation that much heavier.
How it works
The packed result is a single integer array. If the vocabulary is smaller than 65,536, uint16 is enough, and it needs only a quarter of the space of int64. The training loop picks a random position in this array and takes 길이 tokens (the placeholder is the sequence length) to make a batch.
ix = torch.randint(0, len(train) - seq, (batch,), generator=g)
x = torch.stack([train[i:i + seq] for i in ix])
loss = model(x, labels=x).loss # 한 칸 미는 일은 모델 안에서
The MiniMind model takes the same tensor as the input for labels, and inside it aligns logits[..., :-1] with labels[..., 1:] to compute the next-token prediction loss. So there is no need to shift the input and the correct answer separately on the data side.
Validation data must be split off by document. If you cut one array into the first 95% and the last 5%, one document can be split in two and straddle both sides, and if the same document is in both training and validation, the validation loss measures what was memorized. Even after splitting, check once that the same text does not appear verbatim on the training side.
What it looks like in the field
If you train on data with short documents, such as short Q&A and chat logs, with padding, the GPU runs busily but the loss goes down slowly. Utilization is high, but few tokens are actually being learned. If you count the real number of tokens that went into one step at this point, the cause becomes clear right away. Conversely, if after switching to packing the model begins to write unrelated content across document boundaries, first check whether you attached [eos] properly.
The training budget is also counted in tokens. "How many steps" means different things depending on the batch size and length, so you need to compute the tokens one step eats and how many steps make one pass (an epoch) over the whole corpus's tokens, in order to read the bends in the loss curve.
How this course differs from the original MiniMind
MiniMind reads the jsonl with HuggingFace datasets and runs the tokenizer on the spot in __getitem__ for each sample. You do not need to convert a 1.2GB corpus to tokens in advance, but in exchange the CPU keeps cutting tokens while training (which is why the default num_workers is 8). The corpus in this course is under 2MB, so we cut it all at once and save it as a single uint16 array, and the training loop only takes windows from that array. This approach (tokenize in advance into a binary file) is also common in large-scale pretraining — because you do not repeat the same tokenization while going over the data many times, and you can memory-map the file and read only the parts you need.
What you will do in the next lab
You measure document lengths with the reference tokenizer and count how much MiniMind-style padding throws away. You pack the training and validation corpora by document and save them as uint16 arrays, then compute whether validation documents leaked into the training side, what percent of tokens in a packed window can see the previous document, and how many steps one pass takes.