Last Week Had a Better Model. Nobody Can Find It
The Same CSV Differs From Yesterday
Goal
Stamp a fingerprint on the training data to use as its version, check a new batch with a contract that declares what is normal, and measure how much the metric actually moves when the training and evaluation splits overlap.
Why it matters
The place where reproduction breaks is usually not the code but the data. Even if the rows inside change while the name train.csv
stays the same, there is no warning, and only the metrics differ a little. So you attach a hash of the contents to the data as its version, declare the normal range in advance, and test each newly arrived
batch by machine. Quarantining rows that violate the contract instead of silently discarding them is for the same
reason — a silently shrunken sample comes back months later as "why did performance drop?"
The last two steps break one common misconception with a measurement. You would expect the metric to rise noticeably if there is leakage,
but with a model that has weak memory it barely moves.
Steps
- Take fingerprints (hash, size, row count, header) of the three splits.
- Declare a data contract with business rules and confirm that the training data falls within it.
- Test a new batch against the contract and make a list of the violations.
- Quarantine the violating rows and keep the rows to be used for training separately.
- Find the overlapping customers in the broken split.
- Remove the overlap, train twice, and measure the difference in the metric.
- Bundle the facts and limits so far into a data card.
Notes
- Do all the work under
/root/datacontract. First runmkdir -p /root/datacontract. - The materials are in
/opt/fixtures/mlops, and the column descriptions are inDATA-CARD.mdin the same folder. - Row numbers always start from 1 excluding the header. The grader counts by the same rule.
- Common mistakes — reading an empty string as
0and missing a violation, and leaving out the header when rewriting the CSV. - The lab Pod has no volume, so
/rootdisappears when the session ends. If it looks like it will take long, first extend with+시간(the +time button).
Check that a file with the same name is the same file
Create a files object in /root/datacontract/dataset.json and save sha256, bytes, rows (the number of data rows excluding the header), and columns (the list of headers) for each of train.csv, valid.csv, and test.csv. The materials are under /opt/fixtures/mlops.
A file name is not a version. The hash of the contents is the version. Count the rows after skipping the header with the csv module.
Declare what is normal
Save source, source_sha256, row_count, required, ranges, and observed to /root/datacontract/contract.json. source is /opt/fixtures/mlops/train.csv, required is the list of column names in exactly the order written in the header, ranges is the business rules (tenure_months 0..120, monthly_fee 0..200, support_tickets 0..20, late_payments 0..12, usage_hours 0..400, churn 0..1) written as [min, max], and observed is the [min, max] actually observed in train.csv.
A contract is declared from the business, not inferred from the data. But you must confirm that the training data falls within that declaration.
Test the newly arrived batch against the contract
Check /opt/fixtures/mlops/week2.csv with /root/datacontract/contract.json and save source, total_rows, violations, violation_count, and passed to /root/datacontract/validation.json. Each entry of violations has only the three keys row (the row number counting from 1 excluding the header), column, and rule, where rule is required if the value is empty, type if it does not read as a number, and range if it is out of range. Sort by row number ascending, and within the same row by the column order written in the file.
If you gather the check rules in one place in code, you can run them as they are on next week's batch. Be careful not to treat an empty string and 0 as the same thing.
Keep the discarded rows too
Split the rows and save any row that has even one contract violation to /root/datacontract/quarantine.csv and the rest to /root/datacontract/clean_week2.csv. Both files put the same header as the original on the first line and keep the original order.
Rows discarded silently come back next month as "why did the data shrink?" If you have a quarantine file, you can look again later at what was discarded and why.
Check structurally whether the splits overlap
Compare customer_id in /opt/fixtures/mlops/bad_train.csv and /opt/fixtures/mlops/bad_valid.csv and save train and valid (the two file paths), overlap_count, overlap_ratio (the overlap count ÷ the number of evaluation rows, rounded to four decimal places), overlapping_ids (the first 5 in alphabetical order), and split_valid to /root/datacontract/leakage.json.
Count overlap not by rows but by entities. If the same customer is on both sides, the model has already seen the answer for that customer.
Remove the leakage and measure how much the metric moves
From /opt/fixtures/mlops/bad_valid.csv, keep only the rows after removing the customers on the training side and create /root/datacontract/fixed_valid.csv (keep the header and order). Then train twice with bad_train.csv at lr 0.1, epochs 40, seed 7, and save leaky_accuracy (measured with bad_valid.csv), clean_accuracy (measured with fixed_valid.csv), delta (clean − leaky, rounded to four decimal places), threshold (0.05), and visible_in_metric (whether |delta| is at least the threshold) to /root/datacontract/split_effect.json.
Predict the result before you measure. If it differs from your prediction, that difference is what you learn in this step.
Write down what you can say with this data
Save train_sha256, contract_sha256 (the hash of /root/datacontract/contract.json), clean_week2_sha256, violation_count, leakage_overlap_count, approved_for_training, and limitations (40 or more characters) to /root/datacontract/datacard.json. Take the values from the outputs of the earlier steps and match them.
A data card is not a place for boasting but for writing down limits. If you write down what synthetic data cannot say, the next person will not misuse it.