Last Week Had a Better Model. Nobody Can Find It
Nobody Can Find Last Week's Model
Goal
Turn one training session into a unit of record called a "run", record the parameters, metrics, and code and data fingerprints in a ledger, and make it possible to bring that run back to life later. At the end, you confirm from an actual log how a run that left no record disappears.
Why it matters
Building a model well and building that model again are different jobs. Training depends on the random seed, the data version, and the preprocessing code, and if even one of them is not written down, the same number will not come back two months later. What an experiment tracking tool does is not produce fancy graphs but force you to write down enough facts to set one training session up again. Here you design that minimum set yourself — if you fill in by hand the fields that a tool used to fill in automatically, later you will notice right away when you see a ledger with those fields empty.
Steps
- Define the experiment — first write down what to maximize.
- Record the first run in the ledger as one line.
- Change the parameters and increase the runs to four or more.
- Attach code and data fingerprints to every run.
- Pick the best run from the ledger again by the target metric.
- Reproduce that run from the record alone.
- Investigate a run that vanished without a record.
- Export the ledger as a self-contained bundle.
Notes
- Do all the work under
/root/mlops. First runmkdir -p /root/mlops. - The materials are in
/opt/fixtures/mlops. SeeDATA-CARD.mdfor a description of the data. - Example trainer run:
python3 /opt/fixtures/mlops/train_model.py --train /opt/fixtures/mlops/train.csv --valid /opt/fixtures/mlops/valid.csv --lr 0.1 --epochs 40 --seed 7 - The grader reruns the trainer with the arguments written in the ledger and compares the metrics. If you make up numbers, you get caught.
- Common mistakes — writing
matchesorreproducibleas the string"true", and opening the ledger in overwrite mode (w) and erasing the earlier runs. - The lab Pod has no volume. When the session ends,
/rootdisappears entirely, so copy out what you want to keep separately, as in the step 8 bundle.
Write down first what counts as doing well
Save the experiment definition to /root/mlops/experiment.json. name is churn-baseline, objective_metric is valid_accuracy, direction is max, owner is the name of the person responsible (2 or more characters), and dataset is /opt/fixtures/mlops/train.csv.
If you start the experiment before choosing the metric, you will end up choosing a favorable number later. Pin down first what to maximize.
Record the first run in the ledger as one line
Run /opt/fixtures/mlops/train_model.py once and record the result in /root/mlops/runs.jsonl as one line of JSON. One line contains run_id, params (lr, epochs, seed), metrics (valid_accuracy, train_accuracy), data (train and valid paths), and started_at. Copy the metrics from the run output as they are.
The trainer prints one JSON blob to standard output. Do not copy that value by hand; receive it in Python and write it to the ledger, and you cannot make a transcription mistake.
Run four more times and make them comparable
Increase the runs with different combinations of lr, epochs, and seed and leave four or more in /root/mlops/runs.jsonl. run_id must differ for each run, and you must not record the same parameter combination twice. The metrics on every line must equal the values you get when you rerun with those arguments.
If you keep the parameter list in code and loop over it, five lines are made at once. The grader reruns the trainer with each line's arguments and compares the metrics.
Stamp fingerprints on the code and data
Put code_sha256 and data_sha256 on every line of /root/mlops/runs.jsonl. code_sha256 is the SHA-256 hex string of /opt/fixtures/mlops/train_model.py, and data_sha256 is that of /opt/fixtures/mlops/train.csv.
Even with the same parameters, it is a different experiment if the training code or data changes. Hash the file bytes with hashlib.sha256.
Pick the best run from the ledger again
Read the target metric and direction from /root/mlops/experiment.json, pick the best run from /root/mlops/runs.jsonl, and save run_id, metric, value, and selected_by to /root/mlops/best.json. If there is a tie, pick the run that appears first in the ledger.
Pick with code, not by eye. If a person picks, choosing again next week gives a different answer.
Bring that run back to life from the record alone
Read the parameters of the run chosen in /root/mlops/best.json from /root/mlops/runs.jsonl, rerun the trainer, and save the result to /root/mlops/reproduce.json as run_id, params, valid_accuracy, matches, code_sha256, and data_sha256. matches is the boolean true if it equals the ledger's value.
Reproduction is done from "records," not from "memory". Use only the arguments written in the ledger, and the string "true" is not a boolean.
Investigate a run that vanished without a record
/opt/fixtures/mlops/ghost_run.log is a log fragment of a run someone did last week. Save to /root/mlops/incident.json: ghost_metric (the value written in the log), best_recorded_metric (the best value in /root/mlops/runs.jsonl), gap (the difference between the two, rounded to four decimal places), reproducible, missing_fields (the run record fields that cannot be known from the log alone, in alphabetical order), and recovery_plan (40 or more characters).
The only thing readable from the log is a single metric. If you count which of the fields a run record should have are empty, it becomes clear why it cannot be reproduced.
Carry the ledger out before the session ends
Copy runs.jsonl, experiment.json, and best.json under /root/mlops/export/, and save experiment, run_count, best_run_id, and files to /root/mlops/export/manifest.json. Each entry of files holds path (the file name only, without a directory), sha256, and bytes, and best_run_id must be a value recomputed from the ledger it contains.
The lab Pod has no volume, so /root disappears when the session ends. The bundle must stand on its own — a ledger without the objective cannot be interpreted later.