Last Week Had a Better Model. Nobody Can Find It
A Worse Model Went to Production
Goal
Set up a small registry with model versions and aliases (champion and challenger), build a gate that blocks promotion when the metric drops, roll back a promotion that ignored the gate, and finally turn whether the input distribution has moved into a number.
Why it matters
Deploying a model is not copying a file but deciding which version to point at. If production holds the version number directly, rollback becomes a deployment configuration change, but with one layer of alias, rollback becomes changing where the name points. That is why a registry keeps versions, aliases, and history together. A promotion gate adds to this "it goes automatically, but not at just any time" — if you go straight to production as soon as retraining finishes, in some week a worse model quietly goes up. Since cases of a person pushing through after the gate blocked do actually happen, as important as blocking is leaving both the fact that it was blocked and the fact that it was pushed through.
Steps
- Register the model currently in production as version 1.
- Make the champion alias point to version 1.
- Register the retraining result as version 2 and set it up as the challenger.
- Build the promotion gate in code and record the verdict.
- Record the history of the alias moving in an audit log.
- Reproduce a promotion that ignored the gate and record that fact.
- Roll back and write down how much was lost.
- Measure how far the input distribution has shifted and attach it to the retraining decision.
Notes
- Do all the work under
/root/registry. First runmkdir -p /root/registry. - The materials are in
/opt/fixtures/mlops. The trainer's--outwrites the model file for you. - The audit log is append-only. If you edit a line you already wrote, it becomes the current state rather than history.
- Common mistake — deleting version 2 from registry.json when rolling back. Move only the alias.
- The lab Pod has no volume, so
/rootdisappears when the session ends. If it looks like it will exceed 60 minutes, extend in advance with+시간(the +time button).
Register the model currently in production as version 1
Run the trainer with lr 0.1, epochs 40, seed 7 to create a model file, and in /root/registry/registry.json register model as "churn-classifier" and version 1 as the first entry of versions. Each version holds version, run_id, params, metrics, model_path, model_sha256, data_sha256, and created_at. model_sha256 is the hash of the file model_path points to, and data_sha256 is the hash of /opt/fixtures/mlops/train.csv.
The trainer's --out option writes the model file for you. Keep a separate file for each version — if you overwrite, nothing is left to roll back to.
Point the deployment target by name, not number
Save {"champion": 1} to /root/registry/aliases.json. The only aliases used in this lab are champion and challenger.
If you let production look at the version number directly, you have to edit the deployment configuration when you roll back. With one layer of alias, rollback becomes one line in this file.
Register the retraining result as a challenger
Retrain with lr 0.3, epochs 30, seed 3, register version 2 in /root/registry/registry.json, and make challenger in /root/registry/aliases.json point to 2. Version numbers must go up by 1 from 1, and champion must still be 1.
Retraining is not deployment. Set the new version up as a challenger for now, and the gate in the next step judges promotion.
Let a rule, not a person, judge promotion
Write /root/registry/gate.py, read /root/registry/registry.json and /root/registry/aliases.json, and save champion_version, challenger_version, champion_metric, challenger_metric, margin, decision, and reasons to /root/registry/gate.json. margin is 0.005, and decision is promote only when the challenger's valid_accuracy is at least the champion's plus margin, otherwise block. reasons is a list of sentences of 10 or more characters.
The reason for a margin is measurement noise. If you change the champion over a 0.001 difference, you will change it again next week.
Leave a history of the alias moving
Record in /root/registry/audit.jsonl, one line each, the alias settings made so far. Each line holds ts, alias, from (null if it is the first), to, actor, and reason (10 or more characters). If you replay the records from the beginning, you must get exactly the same state as /root/registry/aliases.json.
An audit log holds not "what it is now" but "how it got here". If you write down from, any place that goes out of sync shows up immediately when you replay.
See what a promotion that ignored the gate leaves behind
Reproduce a situation where a person raised champion to version 2 even though the gate blocked it. Change champion in /root/registry/aliases.json to 2, and add to /root/registry/audit.jsonl a line with alias champion, from 1, to 2, writing the verdict of gate.json as it is in that line under the gate_decision key.
Blocking alone is not enough. What the person who pushed it through knew must remain, so that you can talk about the same thing next time.
Roll back and write down how much was lost
Roll champion back to version 1 and record that change in /root/registry/audit.jsonl too (the records must total 4 or more lines and the last line must be champion 2→1). Then save restored_version, bad_version, detected_by (10 or more characters), and metric_delta (version 1's valid_accuracy − version 2's value, rounded to four decimal places) to /root/registry/rollback.json.
A rollback is moving the alias. If you delete the model file or delete the version, what happened disappears with it.
Turn whether the input distribution moved into a number
Using /opt/fixtures/mlops/train.csv as the baseline, measure how far the five features of /opt/fixtures/mlops/week2.csv (tenure_months, monthly_fee, support_tickets, late_payments, usage_hours) have shifted, and save baseline, current, features, max_abs_shift_sigma, alert, and retrain_required to /root/registry/drift.json. Each entry of features holds train_mean, new_mean, pstdev (the population standard deviation of the baseline data), and shift_sigma ((new_mean − train_mean) ÷ pstdev), and empty values or non-numeric values are not counted. alert and retrain_required are whether max_abs_shift_sigma exceeds 1.0.
A distribution shift is not a contract violation — the mean can move even when all values are within the allowed range. That is why you must measure it separately.