TT Lab
Get started
Learn Learning paths Courses

Data Pipelines

Where Did This Number Come From

Continue in TT Lab

One-line summary

Next to each output, write down which inputs (content hashes), which code version, and which parameters it was made with, and prove that record by rerunning with the same inputs and checking whether the bytes are the same.

Why this was needed

The sales in the quarterly report are off by 30 million won from the accounting figure. The pipeline finished normally that day and the log is clean. There is one question — where did this number come from.

When you try to answer, you have little in hand. One output file, and a log line saying "it ran that day." Which drop files went in, whether those files still have the same content as then, whether the code has changed since, what the threshold parameter was at that time — none of it remains. So most investigations go to "let's rerun it." But if the rerun result differs from the one then, you again do not know what changed to make it different.

File names and modification times cannot serve as evidence. It is common for the name to be the same while upstream overwrote the file, and the modification time changes from just copying, moving, or restoring a backup. Conversely, there are cases where the content changed but the time stayed the same. The only thing that can serve as evidence is the hash of the content itself.

Where this course and fde-data diverge

fde-data is about taking apart one chunk of a file a customer gave. That file comes once, and if something looks odd, you can ask the sender. This course is different — a drop falls in the same place with the same name every day, and the same transformation runs repeatedly. So the question is not "how do I read this file" but "which file did that run three months ago read." Lineage becomes a problem only where there is repetition, and reproduction becomes possible because there is repetition.

What to write in the manifest

You leave one JSON with the same name next to the output. There are four things to write.

{
  "code_sha256": "9f2c...",
  "params": {"min_qty": 1},
  "inputs": [{"name": "orders-2026-03-01.csv", "sha256": "4a1e...", "rows": 12}],
  "output": {"name": "shops.csv", "sha256": "7b30...", "rows": 4},
  "run_id": "a41c9d02f7e3b118",
  "created_at": "2026-03-04T02:11:00+00:00"
}

With this one page, all of the earlier questions become answerable. If you rehash the input files now, you can tell whether they are the same as then, if you compare the code hash you can tell whether it changed since, and if you rerun and compare the output hash, you can tell whether the record is correct.

What breaks reproduction

"The same input gives the same output" does not happen by itself. There are four things that quietly break it.

First, the current time. If you write the creation time into the output, the results of running twice are certainly different. If you need the time, write it not in the output but in the manifest, and state firmly that it is a field excluded from the reproduction comparison.

Second, identifiers drawn from random numbers. If you put a run_id drawn with uuid4 on every run into the manifest, the manifest differs even when you run with the same input. You can derive the identifier from the content — if you concatenate the code hash, the parameters, and the output hash and hash them, the same run has the same name and a different run has a different name.

Third, things whose order is not promised. The order of os.listdir is not fixed, and the iteration order of a set varies from run to run even in the same program (Python uses a different seed for string hashing on each run). So you write both the file list and the group list sorted.

Fourth, the order in which floating-point numbers are added. (a + b) + c and a + (b + c) are not the same value in floating point. If the order of lines wavers, the last digit of the total wavers, and then the bytes differ. For values with a fixed number of decimal places, such as amounts, converting to integers (cents) and adding makes this problem disappear entirely.

Dataset level and column level

Lineage has a granularity. The dataset level goes as far as "this table came from those three files." It is easy to build, and it is enough for sweeping the blast radius when an incident happens.

The column level goes down to "this table's amount_cents column came from the input's amount, order_id, and shop." This is the granularity at which you can answer, when upstream tells you they are changing one column, which cell of our outputs will waver.

Column-level lineage must not end with writing it down. As soon as the code changes, that document is stale. The way to prove it is an experiment — you shake one input column, rerun, and see which output columns change. If it changes, there is a dependency, and if it does not, there is none. Measuring this way also reveals "a column we thought we used but do not."

What it looks like in the field

What really matters in practice

What to do in the next lab

You build up the transformer lineage.py, which aggregates the order drops that fall every day, step by step. Starting with digest, which produces a content hash, you confirm why names and times cannot serve as evidence, produce the aggregate output and the manifest together, prove by running twice that the bytes are the same, confirm column-level lineage by experiment, and attach verify, which rebuilds according to the manifest and compares, and trace, which answers "which input that number came from." The grader sets up its own drops each time with different shop names and amounts, actually runs your transformer, and directly shakes the input columns to check that the column lineage you wrote is correct.