TT Lab
Get started
Learn Learning paths Courses

Last Week Had a Better Model. Nobody Can Find It

A Worse Model Went to Production

Continue in TT Lab

Goal

Set up a small registry with model versions and aliases (champion and challenger), build a gate that blocks promotion when the metric drops, roll back a promotion that ignored the gate, and finally turn whether the input distribution has moved into a number.

Why it matters

Deploying a model is not copying a file but deciding which version to point at. If production holds the version number directly, rollback becomes a deployment configuration change, but with one layer of alias, rollback becomes changing where the name points. That is why a registry keeps versions, aliases, and history together. A promotion gate adds to this "it goes automatically, but not at just any time" — if you go straight to production as soon as retraining finishes, in some week a worse model quietly goes up. Since cases of a person pushing through after the gate blocked do actually happen, as important as blocking is leaving both the fact that it was blocked and the fact that it was pushed through.

Steps

  1. Register the model currently in production as version 1.
  2. Make the champion alias point to version 1.
  3. Register the retraining result as version 2 and set it up as the challenger.
  4. Build the promotion gate in code and record the verdict.
  5. Record the history of the alias moving in an audit log.
  6. Reproduce a promotion that ignored the gate and record that fact.
  7. Roll back and write down how much was lost.
  8. Measure how far the input distribution has shifted and attach it to the retraining decision.

Notes

Register the model currently in production as version 1

Run the trainer with lr 0.1, epochs 40, seed 7 to create a model file, and in /root/registry/registry.json register model as "churn-classifier" and version 1 as the first entry of versions. Each version holds version, run_id, params, metrics, model_path, model_sha256, data_sha256, and created_at. model_sha256 is the hash of the file model_path points to, and data_sha256 is the hash of /opt/fixtures/mlops/train.csv.

The trainer's --out option writes the model file for you. Keep a separate file for each version — if you overwrite, nothing is left to roll back to.

Point the deployment target by name, not number

Save {"champion": 1} to /root/registry/aliases.json. The only aliases used in this lab are champion and challenger.

If you let production look at the version number directly, you have to edit the deployment configuration when you roll back. With one layer of alias, rollback becomes one line in this file.

Register the retraining result as a challenger

Retrain with lr 0.3, epochs 30, seed 3, register version 2 in /root/registry/registry.json, and make challenger in /root/registry/aliases.json point to 2. Version numbers must go up by 1 from 1, and champion must still be 1.

Retraining is not deployment. Set the new version up as a challenger for now, and the gate in the next step judges promotion.

Let a rule, not a person, judge promotion

Write /root/registry/gate.py, read /root/registry/registry.json and /root/registry/aliases.json, and save champion_version, challenger_version, champion_metric, challenger_metric, margin, decision, and reasons to /root/registry/gate.json. margin is 0.005, and decision is promote only when the challenger's valid_accuracy is at least the champion's plus margin, otherwise block. reasons is a list of sentences of 10 or more characters.

The reason for a margin is measurement noise. If you change the champion over a 0.001 difference, you will change it again next week.

Leave a history of the alias moving

Record in /root/registry/audit.jsonl, one line each, the alias settings made so far. Each line holds ts, alias, from (null if it is the first), to, actor, and reason (10 or more characters). If you replay the records from the beginning, you must get exactly the same state as /root/registry/aliases.json.

An audit log holds not "what it is now" but "how it got here". If you write down from, any place that goes out of sync shows up immediately when you replay.

See what a promotion that ignored the gate leaves behind

Reproduce a situation where a person raised champion to version 2 even though the gate blocked it. Change champion in /root/registry/aliases.json to 2, and add to /root/registry/audit.jsonl a line with alias champion, from 1, to 2, writing the verdict of gate.json as it is in that line under the gate_decision key.

Blocking alone is not enough. What the person who pushed it through knew must remain, so that you can talk about the same thing next time.

Roll back and write down how much was lost

Roll champion back to version 1 and record that change in /root/registry/audit.jsonl too (the records must total 4 or more lines and the last line must be champion 2→1). Then save restored_version, bad_version, detected_by (10 or more characters), and metric_delta (version 1's valid_accuracy − version 2's value, rounded to four decimal places) to /root/registry/rollback.json.

A rollback is moving the alias. If you delete the model file or delete the version, what happened disappears with it.

Turn whether the input distribution moved into a number

Using /opt/fixtures/mlops/train.csv as the baseline, measure how far the five features of /opt/fixtures/mlops/week2.csv (tenure_months, monthly_fee, support_tickets, late_payments, usage_hours) have shifted, and save baseline, current, features, max_abs_shift_sigma, alert, and retrain_required to /root/registry/drift.json. Each entry of features holds train_mean, new_mean, pstdev (the population standard deviation of the baseline data), and shift_sigma ((new_mean − train_mean) ÷ pstdev), and empty values or non-numeric values are not counted. alert and retrain_required are whether max_abs_shift_sigma exceeds 1.0.

A distribution shift is not a contract violation — the mean can move even when all values are within the allowed range. That is why you must measure it separately.