Rerunning the Past: When a Backfill Doubles Your Numbers
One-line summary
A backfill is not "rerunning a past interval" but "rerunning with the same code as the regular run, without touching the same partition twice, and without overwriting numbers that have already gone out."
Why this was needed
There was a bug. The last two weeks of aggregates were wrong. You rerun that interval with the fixed code. Everyone does up to this point.
The problem comes next. While the backfill runs, the regular run keeps running too. If the two write the same date at the same time, nobody knows which one won. The cumulative aggregate swells by as much as the backfill added. And three days' worth of it had already gone out to customers as a report.
These three are different problems, and they are blocked by different devices. The row-level idempotency dealt with in the earlier lab of this course is only a necessary condition here, not a sufficient one. Even if you have made it so that inserting the same row twice gives the same result, coordination at the interval level must exist separately.
How it works
The first principle is that a backfill must be the same code as the regular run. If you keep a backfill-only script separately, the two codes slowly diverge, and one day nobody can explain why the backfill result differs from the regular result. So the runner only needs to know how to do one thing: "take an interval and recompute the partitions of that interval." A regular run is just the case where that interval is yesterday's single day. Airflow's backfill is also running the same DAG over a past interval, not calling a different DAG.
The second is replacing wholesale at the partition level. If you delete the result of a partition and insert it again, the same answer comes out whether you run it all at once or split it day by day. Without this property, you cannot stop a backfill midway and restart it. And it is usually better to split by partition when running — because when it fails, how far it got is revealed by the partition boundary.
The third is reserving the interval first. The backfill writes in the state store which dates it will use, and the regular run, without touching the dates someone else holds, skips them and reports that fact. Skipping and saying so is better than waiting silently or overwriting silently. This is also the same idea as how, when using locks, stating write intent from the start, as in SQLite's BEGIN IMMEDIATE, is better than failing late.
The fourth is not mixing a backfill into a cumulative aggregate. This is where mistakes happen most often.
누적을 '더하는' 방식
1) 2월 1일부터 14일까지 정기 실행 누적 += 5,182만원 → 5,182만원
2) 2월 9일부터 11일까지 백필 daily 는 제자리에 치환됨
3) 백필 구간을 누적에 더함 누적 += 842만원 → 6,024만원 ← 842만원이 두 번
누적을 '다시 세는' 방식
daily 표 전체를 합쳐 넣는다 누적 = 5,182만원 ← 몇 번 돌려도 같다
Even if the partition table is idempotent, it is useless if the cumulative table is not idempotent. Do not add to the cumulative total; derive it. If you recount from the source partitions, the value does not waver no matter how many times you run the backfill.
The fifth is that there are places you must not go back to. A report already sent, an alert already issued, an amount already settled are not data but events. You seal that interval, and when a new value comes out, instead of overwriting, you record it separately as a correction. Only then can you say both "what did we send at that time" and "what is correct now."
What it looks like in the field
First, building a separate backfill script. A script made in a hurry to use once remains and is still running half a year later. Exception handling that only that script knows appears, and the answers of the two paths diverge.
Second, putting the whole interval in one transaction. If you commit 14 days at once, when it fails on the 13th day you have to start over from the beginning, and in the meantime the state store knows nothing. If you finish per partition, a restart point comes for free.
Third, believing "now is not a time when the regular run runs" without a reservation. That belief breaks because of retries, manual runs, and time differences. A reservation is a promise that code keeps, and a time window is a promise that people keep.
Fourth, leaving the backfill range only in logs. If what was rerun, when, and by whom is not left in the state store, there is no basis to retrace when the numbers look odd. And do not estimate whether something was a backfill from its run time — even on the same machine, run time can swing by up to a factor of two.
What really matters in practice
- A backfill and a regular run are the same function. The only difference should be the interval argument.
- Replace at the partition level, and commit per partition. A restart point arises.
- Reserve the interval, and skip other people's intervals and report. Do not wait or overwrite.
- Derive the cumulative total; do not add to it. An additive cumulative total is certainly wrong when it meets a backfill.
- Seal what has already gone out and leave corrections. If you overwrite, the fact of that time disappears.
What to do in the next lab
After creating fourteen days of order partitions, you build up the runner runner.py step by step. You make it take an interval and replace at the partition level, confirm that the answer from running everything at once and the answer from running day by day are the same, and write in numbers how far an additive cumulative total and a derived cumulative total diverge after a backfill. Then you attach interval reservation so that the regular run skips other people's intervals and reports, and finally you seal the intervals already sent out and, for a voucher that arrived late, leave a correction instead of overwriting. The grader sets up its own partitions each time with different dates and amounts, actually runs your runner, and reads the state store directly to compare.