Last Week Had a Better Model. Nobody Can Find It
A Name Is Not a Version
In one line
Version data not by name but by a hash of its contents, and declare the normal range from the business side rather than inferring it from the data, then test each newly collected batch against it by machine.
Why this was needed
The place where reproduction breaks is usually not the code but the data. Git protects the training script, but nobody protects train.csv. One day, if the upstream pipeline
loads duplicates once, the rows grow, and next week if someone changes the missing-value handling, the values change. The file name
stays the same, training finishes normally, and the metrics move a little. This combination hides the longest.
The second problem is that the definition of "normal" is written nowhere. Even if a row with a negative monthly fee comes in, Python reads it fine as a float, and the model trains on that value. People notice only weeks later when predictions start to look strange, and by then the model trained on that data is already in production.
How it works
Data versions are captured with fingerprints. If you compute the SHA-256 of the file bytes, you can answer in one step whether a file with the same name is the same file (hashlib). MLflow's dataset object carries the name, digest, source location, schema, and profile together, following the same idea (MLflow Dataset Tracking). If the digest is attached to the run record, "what data was this model built from" becomes a lookup rather than a guess.
The normal range is declared as a schema. The TensorFlow Data Validation documentation describes a schema as "a description in code of the properties the input data must satisfy," and uses an approach of comparing statistics against the schema to find anomalies. The problems the same document lists as common are missing values, a label mixed in as a feature so that the model sees the correct answer in advance, and values outside the expected range (TensorFlow Data Validation).
계약(선언) monthly_fee 0~200, late_payments 필수
관찰(계산) train.csv 의 monthly_fee 는 9.3 ~ 88.69
검증(대조) week2.csv 240행 중 4행이 계약 위반
The validation rules themselves take only a few lines. Is a required column empty, does it read as a number, does it fall within the declared range. What is hard is not the rules but deciding in advance what to do when you meet a violation. Whether to stop entirely, to skip only the violating rows, or to leave a warning and train as is can differ by column.
It matters here to separate declaration from observation. If you take today's minimum and maximum directly as the rule, an outlier that happened to come in today becomes tomorrow's rule. The business decides the range, and the data is only checked for whether it falls inside that range.
Split leakage is another axis. The scikit-learn documentation defines leakage as "information that cannot be used at prediction time being used to build the model," and explains that as a result the performance estimate becomes overly optimistic and is worse on real new data. It names as the most common cause failing to properly separate the training and evaluation subsets (Common pitfalls).
What it looks like in the field
In the first few days after adding contract validation, alerts pour in. Most of them are not because the data is bad but because the contract was written narrower than reality. If you turn validation off at that point, you are back to square one, so it is better to read the alerts one by one and fix the contract. That process becomes the first time the team agrees on "what shape our data has".
The second is the habit of silently discarding. If you skip rows that violate the contract with try/except, training
keeps running but the fact that the sample shrank is recorded nowhere. Months later it comes back as "why did performance
drop?" If you keep the discarded rows in a quarantine file, you can retrace the reasons later.
The third is an attempt to catch leakage with a metric. With a model with a large memory, leakage raises the metric a lot, but with a simple linear model it barely moves. So a split must not be suspected by looking at the numbers but checked by whether identifiers overlap.
The fourth is putting off the data card. If you do not write down where this data came from and what it cannot say, the next person uses it without knowing those limits. That is how accidents happen, such as attaching a model built on synthetic data directly to real customer handling. A card does not need to be long — the source, the fingerprint, the split rule, and a couple of lines saying "you cannot say this kind of thing with this" are enough, and those couple of lines save one meeting months later.
What you will do in the next lab
You take fingerprints of the three splits, declare a data contract with business rules, test a new batch against it, and quarantine the violating rows. Then you find a broken split where the same customer is in both training and evaluation, remove the overlap, and measure yourself how much the metric actually moves.