Last Week Had a Better Model. Nobody Can Find It
Blocking It and Recording It
In one line
Retraining can run automatically, but promotion must pass a gate, and if a person pushed through something the gate blocked, that fact must also remain in the record.
Why this was needed
Once you build a retraining pipeline, one temptation appears: since training has finished, to carry it straight on into deployment. In most weeks nothing happens. The problem is that nothing happening does not mean it succeeded. Even in a week when upstream data broke once, the pipeline lights up the same green, and a worse model quietly goes up to production.
The MLOps maturity document compiled by Google draws a clear line at this point. Automating the training pipeline (continuous training) and putting its output into production (continuous delivery) are different stages, and the latter must have a model validation step in it (MLOps: Continuous delivery and automation pipelines in machine learning).
How it works
A gate is a few lines of rules, simpler than you might think. What is hard is not writing the rules but whether the numbers the rules read are prepared in a trustworthy way. That is why the earlier modules come first.
입력 champion 별칭이 가리키는 버전의 평가 지표
challenger 별칭이 가리키는 버전의 평가 지표
데이터 계약 검증 결과
규칙 challenger >= champion + margin 이고 계약 검증 통과
출력 promote | block + 왜 그렇게 판단했는지
The reason for having a margin is measurement noise. Even on the same data, with a different seed accuracy wobbles at the third decimal place. If the margin is 0, that wobble alone changes the champion, and it changes again next week. Deployments do not just become more frequent; they become meaninglessly frequent.
MLflow gives, as an example in its documentation, attaching the result of this judgment as a version tag. You attach validation_status: pending to a version
awaiting validation and approved to one that passed
(MLflow Model Registry).
If you leave the gate's verdict as a tag, automation rather than a person can pick the next step based on that condition.
You also need to decide where the numbers the gate reads come from. If you draw the evaluation split anew each time, the metric wobbles and the gate becomes meaningless, and conversely, if you keep using only a fixed split, a model overfit to that split passes the gate. So you usually take the fixed split as the standard but add a new split periodically and look at both numbers together.
And rollback. The reason you put aliases in the earlier module pays off here. A rollback is
changing the version champion points to, and that change becomes one line in the audit log.
If the record also carries the previous value (from), replaying it from the beginning must produce the current state,
and if it does not, it means there was an unrecorded change somewhere. This one property
turns "who changed what and when" from a guess into a check you can verify by recomputing.
The gate's verdict must be readable by a person. If only the single word block remains, the next
person cannot tell why it was blocked and comes to doubt the gate, and a gate that is doubted is soon turned off.
So next to the verdict you leave, as sentences, the two numbers that were compared, the margin, and the condition under which it was caught.
What it looks like in the field
Once you build a gate, bypass requests inevitably come. The demo is tomorrow, the customer is waiting, and they say this one is certain. At that point, instead of turning the gate off, it lasts longer to design it so that it can be overridden but the fact of overriding is recorded. If you try to block what you cannot block, people build a new path outside the gate, and then not even a record remains.
The second most common is when the gate looks at only one metric. If you look only at overall accuracy, a model that got much worse on a small group passes. Adding metrics is easy, but for each added metric you must decide "how much worse before we block," so it needs agreement. Writing that agreement down in code is the gate's real value.
The third is silence after a rollback. They do the rollback well, but nobody writes down why that version was bad, so two months later they repeat the same mistake with similar parameters.
The fourth is observation after passing the gate. Promotion is not the end but the beginning, so you need a window of a few days to watch how the new champion actually behaves in production. That is why many teams first send the challenger only part of the traffic and then move the alias. Here too, the basis for the judgment is written in the same place — if the numbers the gate saw and the numbers seen during the observation period are scattered in different places, then in the meeting that decides whether to roll back, two people show up with different tables.
What you will do in the next lab
You set up a small registry with versions and aliases, register the retraining result as a challenger, and build a promotion gate with a margin in code. Then you reproduce a situation where a person pushed through a promotion the gate blocked, and after rolling back, you replay the audit log from the beginning and check that it matches the current state. Finally, you turn how far the input distribution has moved into a number.