TT Lab
Get started
Learn Learning paths Courses

MiniMind — Train a Small Language Model Yourself, End to End

Measure the compute padding wastes and pack the corpus

Continue in TT Lab

Goal

Cut the pretraining corpus with the reference tokenizer and measure what percent of computation MiniMind-style padding throws away, and make the training and validation data as uint16 arrays (packing) with documents wrapped in [bos] … [eos] and joined together. You also check validation leakage, document boundaries, and the training budget in numbers.

Why it matters

A matrix multiplication does not know what is padding. Padding slots are only left out of the loss, and the computation is done as it is. If documents are short and you set the maximum length long, the GPU is busy but few tokens are learned. Packing removes that waste, but in exchange it puts several documents in one window so that the later document sees the earlier one — the tradeoff most pretraining accepts. Mistakes made in the data preparation stage show up only after training ends. If validation documents are mixed into the training side, the validation loss lies, and if you do not know the number of tokens in a step, you cannot read the loss curve. So you check them in numbers before training.

Steps

  1. With /opt/mm/ref/tokenizer.json, count the tokens of each document in /opt/mm/data/pretrain.jsonl and write them to /root/mm/data/lengths.json as docs, mean, max, and p95.
  2. Like MiniMind's PretrainDataset, if you wrap one document as [bos] + 토큰(최대 126) + [eos] (the Korean word in it stands for the tokens, at most 126) and pad it to length 128, write how many padding tokens there are to /root/mm/data/padding.json as max_len, rows, pad_tokens, pad_fraction, and truncated_docs.
  3. Save a uint16 array, made by wrapping each document of the training corpus as [bos](1) 토큰들 [eos](2) (the Korean word in it stands for the document's tokens) and joining them, to /root/mm/data/train.npy.
  4. Save /opt/mm/data/pretrain_val.jsonl to /root/mm/data/val.npy in the same way.
  5. Write the number of validation documents that appear verbatim in the training corpus to /root/mm/data/leak.json as val_docs and dup_in_train.
  6. In windows cut from train.npy at 128 tokens each with no overlap, write the average number of documents per window and the fraction of tokens that can see the previous document to /root/mm/data/windows.json as seq_len, windows, docs_per_window, and cross_doc_fraction.
  7. For batch 16 and length 128, write the number of tokens per step and the number of steps per pass (epoch) to /root/mm/data/budget.json as train_tokens, tokens_per_step, and steps_per_epoch.
  8. In /root/mm/data/report.md, write the three sections ## 패딩, ## 패킹, and ## 검증 분리 (the Korean headings mean "Padding", "Packing", and "Validation split"), and include the pad_fraction from step 2 and the cross_doc_fraction from step 6 as numbers.

Notes

How many tokens is a document

With /opt/mm/ref/tokenizer.json, cut the text of /opt/mm/data/pretrain.jsonl document by document, and write the document count, mean, maximum, and 95th percentile (numpy.percentile(길이, 95), where the placeholder is the list of lengths) to /root/mm/data/lengths.json as docs, mean, max, and p95.

From the mean and maximum alone, you can guess how much padding there will be. If you set the maximum length to the maximum, nothing gets cut, but a document of average length fills all the rest with padding.

The share MiniMind-style padding throws away

Assuming you wrap one document as [bos] + 토큰(126개까지 자름) + [eos] (the Korean text in it means the tokens, cut to at most 126) and pad to length 128, write the padding token count, its ratio to the total (document count × 128), and the number of truncated documents to /root/mm/data/padding.json as max_len, rows, pad_tokens, pad_fraction, and truncated_docs.

MiniMind's PretrainDataset.__getitem__ does exactly this calculation — it cuts to max_length - 2, attaches bos and eos, and fills the rest with pad. Padding is left out of the loss (-100), but the matrix multiplication goes through those slots too.

Join the training corpus into one line

Write the script /root/mm/data/pack.py (arguments: input jsonl, .npy to save), wrap the documents of /opt/mm/data/pretrain.jsonl in file order as 1(bos) 토큰들 2(eos) (the Korean word in it stands for the document's tokens) and join them, and save it as a uint16 array to /root/mm/data/train.npy.

If you change the file order or shuffle, it will differ from the array the grader rebuilds — shuffling is done when the training loop draws batches. Since the vocabulary is 1024, uint16 (max 65535) is enough, and it is a quarter the size of int64.

The validation corpus the same way

Pack /opt/mm/data/pretrain_val.jsonl in the same way as step 3 and save it to /root/mm/data/val.npy.

You must make the validation data in the same way as training to compare the two losses. If you make the step 3 script take the input and output paths as arguments, it is one line.

Did validation documents leak

Write the number of validation documents (the text of pretrain_val.jsonl) that appear verbatim in the training corpus (pretrain.jsonl) to /root/mm/data/leak.json as val_docs and dup_in_train.

If you make the training-side text into a set, you only need to look up each one once. If it is not 0, the validation loss of those documents measures what was memorized.

How many documents per window, and what percent of tokens cross the boundary

Cut train.npy into 128-token windows with no overlap (discard the leftover tail), and write the average [bos] count per window (docs_per_window) and the fraction of tokens that can see the previous document (cross_doc_fraction) to /root/mm/data/windows.json together with seq_len and windows. A token "can see the previous document" means that within the window, a document start ([bos]) has occurred at least once up to that token's position (including itself) — however, if the window's first token is [bos], that document is the first to start in the window, so it is not counted.

Use numpy.cumsum(창 == 1, axis=1) (the placeholder is the windows array) to get the number of bos seen so far at each position, and subtract 1 for windows whose first position is bos. Since the causal mask lets each token see all preceding tokens, these tokens see the previous document unless you build a per-document mask separately.

How many steps is one pass

Write the number of tokens one step eats when training with batch 16 and length 128, and the number of steps needed to see the whole of train.npy once (rounded up), to /root/mm/data/budget.json as train_tokens, tokens_per_step, and steps_per_epoch.

When you train hundreds of steps in the next module, this number tells you how many passes over the corpus that is. If you see the same data for several passes, the training loss starts to fall faster than the validation loss.

Leave the reasons for how you prepared the data

In /root/mm/data/report.md, write the three sections ## 패딩, ## 패킹, and ## 검증 분리 (the Korean headings mean "Padding", "Packing", and "Validation split"), and include the pad_fraction from step 2 and the cross_doc_fraction from step 6 as numbers (decimals or percentages).

Write one line each on which of padding and packing you would choose and what you accept in exchange. In the validation split section, you can put the result of step 5.