Where Did This Number Come From
One-line summary
Next to each output, write down which inputs (content hashes), which code version, and which parameters it was made with, and prove that record by rerunning with the same inputs and checking whether the bytes are the same.
Why this was needed
The sales in the quarterly report are off by 30 million won from the accounting figure. The pipeline finished normally that day and the log is clean. There is one question — where did this number come from.
When you try to answer, you have little in hand. One output file, and a log line saying "it ran that day." Which drop files went in, whether those files still have the same content as then, whether the code has changed since, what the threshold parameter was at that time — none of it remains. So most investigations go to "let's rerun it." But if the rerun result differs from the one then, you again do not know what changed to make it different.
File names and modification times cannot serve as evidence. It is common for the name to be the same while upstream overwrote the file, and the modification time changes from just copying, moving, or restoring a backup. Conversely, there are cases where the content changed but the time stayed the same. The only thing that can serve as evidence is the hash of the content itself.
Where this course and fde-data diverge
fde-data is about taking apart one chunk of a file a customer gave. That file comes once, and if something looks odd, you can ask the sender. This course is different — a drop falls in the same place with the same name every day, and the same transformation runs repeatedly. So the question is not "how do I read this file" but "which file did that run three months ago read." Lineage becomes a problem only where there is repetition, and reproduction becomes possible because there is repetition.
What to write in the manifest
You leave one JSON with the same name next to the output. There are four things to write.
- Inputs: the file name, content hash, byte count, and record count. Writing only the name is useless.
- Code: the hash of the transformer's source (or the commit). A human-written name like "version 1.2" lies.
- Parameters: all the values given to that run. Write the defaults too — defaults change later.
- Output: the content hash and record count of the output.
{
"code_sha256": "9f2c...",
"params": {"min_qty": 1},
"inputs": [{"name": "orders-2026-03-01.csv", "sha256": "4a1e...", "rows": 12}],
"output": {"name": "shops.csv", "sha256": "7b30...", "rows": 4},
"run_id": "a41c9d02f7e3b118",
"created_at": "2026-03-04T02:11:00+00:00"
}
With this one page, all of the earlier questions become answerable. If you rehash the input files now, you can tell whether they are the same as then, if you compare the code hash you can tell whether it changed since, and if you rerun and compare the output hash, you can tell whether the record is correct.
What breaks reproduction
"The same input gives the same output" does not happen by itself. There are four things that quietly break it.
First, the current time. If you write the creation time into the output, the results of running twice are certainly different. If you need the time, write it not in the output but in the manifest, and state firmly that it is a field excluded from the reproduction comparison.
Second, identifiers drawn from random numbers. If you put a run_id drawn with uuid4 on every run into the manifest, the manifest differs even when you run with the same input. You can derive the identifier from the content — if you concatenate the code hash, the parameters, and the output hash and hash them, the same run has the same name and a different run has a different name.
Third, things whose order is not promised. The order of os.listdir is not fixed, and the iteration order of a set varies from run to run even in the same program (Python uses a different seed for string hashing on each run). So you write both the file list and the group list sorted.
Fourth, the order in which floating-point numbers are added. (a + b) + c and a + (b + c) are not the same value in floating point. If the order of lines wavers, the last digit of the total wavers, and then the bytes differ. For values with a fixed number of decimal places, such as amounts, converting to integers (cents) and adding makes this problem disappear entirely.
Dataset level and column level
Lineage has a granularity. The dataset level goes as far as "this table came from those three files." It is easy to build, and it is enough for sweeping the blast radius when an incident happens.
The column level goes down to "this table's amount_cents column came from the input's amount, order_id, and shop." This is the granularity at which you can answer, when upstream tells you they are changing one column, which cell of our outputs will waver.
Column-level lineage must not end with writing it down. As soon as the code changes, that document is stale. The way to prove it is an experiment — you shake one input column, rerun, and see which output columns change. If it changes, there is a dependency, and if it does not, there is none. Measuring this way also reveals "a column we thought we used but do not."
What it looks like in the field
- Same name, different content. Upstream fixed yesterday's file and uploaded it again, but the name was the same, so nobody knew. If you had left a content hash, the next run would catch it immediately.
- "We didn't change the code." This is the case where a library version went up or a default parameter changed. If you write the code hash and the parameters together, this conversation ends in 30 seconds.
- A reproduction script that does not reproduce. You reran it to investigate and got a different answer all three times. It is usually one of the four above.
- Mismatch between the lineage document and the code. The lineage written on the wiki is from half a year ago. There is no way to notice this unless you measure again by experiment.
What really matters in practice
- Create the output and the manifest together. If you attach it later, it is certainly missed.
- Judge sameness by content hash, not by name and time.
- State reproduction not as a claim but as the result of running twice and comparing the bytes.
- Gather the field that breaks reproduction (the creation time) into one, and write down that this field is excluded from the comparison.
- Remeasure column lineage by experiment, not by document.
What to do in the next lab
You build up the transformer lineage.py, which aggregates the order drops that fall every day, step by step. Starting with digest, which produces a content hash, you confirm why names and times cannot serve as evidence, produce the aggregate output and the manifest together, prove by running twice that the bytes are the same, confirm column-level lineage by experiment, and attach verify, which rebuilds according to the manifest and compares, and trace, which answers "which input that number came from." The grader sets up its own drops each time with different shop names and amounts, actually runs your transformer, and directly shakes the input columns to check that the column lineage you wrote is correct.