TT Lab
Get started
Learn Learning paths Courses

Last Week Had a Better Model. Nobody Can Find It

A Distribution That Moves Inside the Range

Continue in TT Lab

In one line

Drift is not a contract violation — the distribution can shift even when all values are within the allowed range, and catching it requires not a range check but a metric that compares distributions.

Why this was needed

The data contract built in the earlier module is powerful, but there is one thing it cannot catch. Even if the mean usage time moves from 38 hours to 55 hours, those values are still within the range "0–400". The contract passes, training is normal, and there are no alerts. But the relationships the model learned were built on the old distribution, so its predictions begin to miss little by little.

The problem of labels not arriving right away is layered on top. For churn prediction, you learn whether that customer really churned only a month later. To notice an anomaly through accuracy you have to wait a month, but the input distribution can be measured right now. So drift monitoring is an early signal used while waiting for labels.

How it works

The TFX documentation explains data validation in three parts: schema-based validation, detection of training-serving skew, and finding drift by looking across a series of data (TensorFlow Data Validation). The third is the subject of this module, and the point is that it is a separate check from the first two.

The simplest form is to remember the mean and standard deviation of the baseline data and look at how many sigmas the mean of the new data has shifted. The population standard deviation can be computed directly with the standard library (statistics).

shift_sigma = (새 데이터의 평균 − 기준의 평균) ÷ 기준의 모집단 표준편차

|shift_sigma| 0.1 정도   잡음 범위
|shift_sigma| 1.0 이상   기준 분포에서 확실히 벗어남 → 경보

A single sigma cannot measure everything. A categorical feature has no mean, and for a value with a distribution heavily skewed to one side, a mean shift looks smaller than the real change. So in practice people also use an approach of splitting into bins and measuring the difference in proportions, or comparing quantiles. Whichever way you choose, the common point is the same — the baseline distribution must be stored somewhere, and that stored artifact must have a version like a model.

What to use as the baseline is the next decision. If you fix it to the data used for training, you measure the distance from "the world the model learned," and if you use last week's production data, you measure "did it suddenly change." They are different questions, so usually you keep both.

Once you have turned it into a metric, the next part is the monitoring system's job. You emit one value per feature and raise an alert when it crosses a threshold. It is easier to read later if metric names follow the convention of attaching the unit as a suffix and using base units (Metric and label naming). You can also connect the alert directly to a retraining trigger, but a promotion gate must sit in front of it — there is no guarantee that a model retrained because the distribution changed will be better.

Measuring drift does not mean you must change the model right away. The first purpose of monitoring is not to be surprised. If you know the distribution is moving, then when accuracy drops a month later you do not have to look for the cause from scratch. Conversely, if you are measuring nothing, that whole month becomes guesswork.

What it looks like in the field

The most common failure is too many alerts. If you put a threshold on each of 200 features, a few cross it every day. After a few days nobody looks. So in practice you monitor only the few top features the model depends on heavily, or fold the per-feature values into one summary value and alert on that.

The second is reading drift straight away as an incident. When you overhaul the pricing plans, the usage time distribution naturally changes. That is not the data breaking but the world changing, and the needed response is retraining, not rollback. Conversely, if the upstream pipeline has switched the unit from minutes to hours, what is needed is a fix, not retraining. The metric does not distinguish the two, so an alert must always carry "what to check".

The third is connecting an alert straight to automatic retraining. There is no guarantee anywhere that a model retrained because the distribution moved will be better, and it might even train on broken data as it is. So between the alert and retraining there must be data contract validation, and between retraining and deployment there must be a promotion gate.

The fourth is not updating the baseline. If you changed the champion by retraining, the drift baseline must also move to the new training data, and if you miss this, the alert stays on forever.

The fourth is concluding the cause from a single feature. Whether the rise in usage time is really a change in user behavior, a change in how the data is collected, or a mass influx of a particular customer group is not separated by that one number alone. So next to an alert there must always be the sample count and the upper quantiles for the same period. Looking only at the mean, you cannot tell the sample growing tenfold from the values shifting.

What to check in the next quiz

You check what contract validation and drift detection each catch, why you measure the input distribution when labels arrive late, and what more is needed when you connect a drift alert directly to retraining.