MiniMind — Train a Small Language Model Yourself, End to End
Measure the compute padding wastes and pack the corpus
Goal
Cut the pretraining corpus with the reference tokenizer and measure what percent of computation MiniMind-style padding throws away, and make the training and validation data as uint16 arrays (packing) with documents wrapped in [bos] … [eos] and joined together. You also check validation leakage, document boundaries, and the training budget in numbers.
Why it matters
A matrix multiplication does not know what is padding. Padding slots are only left out of the loss, and the computation is done as it is. If documents are short and you set the maximum length long, the GPU is busy but few tokens are learned. Packing removes that waste, but in exchange it puts several documents in one window so that the later document sees the earlier one — the tradeoff most pretraining accepts. Mistakes made in the data preparation stage show up only after training ends. If validation documents are mixed into the training side, the validation loss lies, and if you do not know the number of tokens in a step, you cannot read the loss curve. So you check them in numbers before training.
Steps
- With
/opt/mm/ref/tokenizer.json, count the tokens of each document in/opt/mm/data/pretrain.jsonland write them to /root/mm/data/lengths.json asdocs,mean,max, andp95. - Like MiniMind's
PretrainDataset, if you wrap one document as[bos] + 토큰(최대 126) + [eos](the Korean word in it stands for the tokens, at most 126) and pad it to length 128, write how many padding tokens there are to /root/mm/data/padding.json asmax_len,rows,pad_tokens,pad_fraction, andtruncated_docs. - Save a
uint16array, made by wrapping each document of the training corpus as[bos](1) 토큰들 [eos](2)(the Korean word in it stands for the document's tokens) and joining them, to /root/mm/data/train.npy. - Save
/opt/mm/data/pretrain_val.jsonlto /root/mm/data/val.npy in the same way. - Write the number of validation documents that appear verbatim in the training corpus to /root/mm/data/leak.json as
val_docsanddup_in_train. - In windows cut from
train.npyat 128 tokens each with no overlap, write the average number of documents per window and the fraction of tokens that can see the previous document to /root/mm/data/windows.json asseq_len,windows,docs_per_window, andcross_doc_fraction. - For batch 16 and length 128, write the number of tokens per step and the number of steps per pass (epoch) to /root/mm/data/budget.json as
train_tokens,tokens_per_step, andsteps_per_epoch. - In /root/mm/data/report.md, write the three sections
## 패딩,## 패킹, and## 검증 분리(the Korean headings mean "Padding", "Packing", and "Validation split"), and include the pad_fraction from step 2 and the cross_doc_fraction from step 6 as numbers.
Notes
- Cut with
tok.encode(글, add_special_tokens=False).ids(the placeholder is the text), and save withnumpy.save(경로, numpy.array(ids, dtype=numpy.uint16))(the placeholder is the file path). - Common mistakes: forgetting
[bos]and[eos]before joining documents (the length differs from the array the grader rebuilds), saving as int64, and counting even the case where a window's first token is[bos]as "crossed the boundary". - Sources: MiniMind lm_dataset.py — PretrainDataset · numpy.save
How many tokens is a document
With /opt/mm/ref/tokenizer.json, cut the text of /opt/mm/data/pretrain.jsonl document by document, and write the document count, mean, maximum, and 95th percentile (numpy.percentile(길이, 95), where the placeholder is the list of lengths) to /root/mm/data/lengths.json as docs, mean, max, and p95.
From the mean and maximum alone, you can guess how much padding there will be. If you set the maximum length to the maximum, nothing gets cut, but a document of average length fills all the rest with padding.
The share MiniMind-style padding throws away
Assuming you wrap one document as [bos] + 토큰(126개까지 자름) + [eos] (the Korean text in it means the tokens, cut to at most 126) and pad to length 128, write the padding token count, its ratio to the total (document count × 128), and the number of truncated documents to /root/mm/data/padding.json as max_len, rows, pad_tokens, pad_fraction, and truncated_docs.
MiniMind's PretrainDataset.__getitem__ does exactly this calculation — it cuts to max_length - 2, attaches bos and eos, and fills the rest with pad. Padding is left out of the loss (-100), but the matrix multiplication goes through those slots too.
Join the training corpus into one line
Write the script /root/mm/data/pack.py (arguments: input jsonl, .npy to save), wrap the documents of /opt/mm/data/pretrain.jsonl in file order as 1(bos) 토큰들 2(eos) (the Korean word in it stands for the document's tokens) and join them, and save it as a uint16 array to /root/mm/data/train.npy.
If you change the file order or shuffle, it will differ from the array the grader rebuilds — shuffling is done when the training loop draws batches. Since the vocabulary is 1024, uint16 (max 65535) is enough, and it is a quarter the size of int64.
The validation corpus the same way
Pack /opt/mm/data/pretrain_val.jsonl in the same way as step 3 and save it to /root/mm/data/val.npy.
You must make the validation data in the same way as training to compare the two losses. If you make the step 3 script take the input and output paths as arguments, it is one line.
Did validation documents leak
Write the number of validation documents (the text of pretrain_val.jsonl) that appear verbatim in the training corpus (pretrain.jsonl) to /root/mm/data/leak.json as val_docs and dup_in_train.
If you make the training-side text into a set, you only need to look up each one once. If it is not 0, the validation loss of those documents measures what was memorized.
How many documents per window, and what percent of tokens cross the boundary
Cut train.npy into 128-token windows with no overlap (discard the leftover tail), and write the average [bos] count per window (docs_per_window) and the fraction of tokens that can see the previous document (cross_doc_fraction) to /root/mm/data/windows.json together with seq_len and windows. A token "can see the previous document" means that within the window, a document start ([bos]) has occurred at least once up to that token's position (including itself) — however, if the window's first token is [bos], that document is the first to start in the window, so it is not counted.
Use numpy.cumsum(창 == 1, axis=1) (the placeholder is the windows array) to get the number of bos seen so far at each position, and subtract 1 for windows whose first position is bos. Since the causal mask lets each token see all preceding tokens, these tokens see the previous document unless you build a per-document mask separately.
How many steps is one pass
Write the number of tokens one step eats when training with batch 16 and length 128, and the number of steps needed to see the whole of train.npy once (rounded up), to /root/mm/data/budget.json as train_tokens, tokens_per_step, and steps_per_epoch.
When you train hundreds of steps in the next module, this number tells you how many passes over the corpus that is. If you see the same data for several passes, the training loss starts to fall faster than the validation loss.
Leave the reasons for how you prepared the data
In /root/mm/data/report.md, write the three sections ## 패딩, ## 패킹, and ## 검증 분리 (the Korean headings mean "Padding", "Packing", and "Validation split"), and include the pad_fraction from step 2 and the cross_doc_fraction from step 6 as numbers (decimals or percentages).
Write one line each on which of padding and packing you would choose and what you accept in exchange. In the validation split section, you can put the result of step 5.