The Decision to Stop Is Only Cheap Before You Start
In one line
You cannot decide during the run whether to stop on an anomaly — by then the time already spent and the half-changed state push you toward continuing, so the place to stop has to be written down on paper before the run.
Why decide in advance
Suppose you have applied a change halfway and the control group numbers come out different from the plan. The only judgment you need now is "Do we stop?", but the conditions of the person making that judgment are completely different from 30 minutes ago.
You have already spent two hours, the customer's owner is sitting next to the screen, and the data is only half changed. In this state people lean toward continuing almost every time. "We've come this far, let's finish and see", "If we stop now the state gets even weirder". Both sentences sound reasonable at that moment, and when you reread them later in an incident report, both are excuses.
That is why you decide the place to stop while nothing has changed yet, that is, while stopping costs nothing. The me who is running is not the person who rewrites that sentence but the person who keeps it.
How it works
Usable abort criteria must contain all three of the following.
관측할 수 있는 지표 무엇을 보고 판단하는가. 느낌이 아니라 세어지는 것.
넘으면 안 되는 값 그 지표가 얼마가 되면 이상인가. 부등호와 숫자로.
그때 할 행동 멈추고 무엇을 하는가. 되돌린다 / 보류하고 보고한다.
If even one of the three is missing, room for interpretation appears during the run, and room for interpretation is always used in favor of continuing. "Stop if something looks wrong" is not an abort criterion but a resolution.
A partially applied state is the most dangerous. The state before the change and the state after the change can both be explained, but a half-changed state is in no document. So stopping usually means not "standing there" but "rolling back". If you cannot roll back and have to stand there, announcing that very fact immediately is part of the abort action.
Time is a criterion too. If the change window is 60 minutes, the apply must not end at 60 minutes; it must end with the time needed for the rollback and the time needed to check after the rollback subtracted. If you run past the window with no time left to roll back, then even with abort criteria in hand, you have no means to execute them.
And the best method is not to create an irreversible change at all. If you leave a deletion marker instead of deleting and delete after a grace period, one irreversible step is split into two reversible steps. Stopping writes first and watching one cycle before dropping a column is the same approach. Asking whether you can split the change this way comes before writing good abort criteria.
What it looks like in the field
It often happens that right before the run, the requester asks you to change the condition in just one place. It sounds as if a single phrase, "Oh, leave that order out", is enough, but once the condition changes, the scoping, the dry run, and the backup have all been done on a different condition. The right response is to reflect that one place and go through the earlier steps again, and if there is no time, to run with the original condition and handle the difference as a separate change. Fixing only the condition on the spot and running it is the most common path to an incident.
And one more thing: after you actually carry out an abort, the fact that you stopped is itself something to report. If you pass over it saying nothing happened because you rolled back, the customer will later find the trace in the logs and read it as us hiding it. Writing up and sending, within the same day, the reason you stopped, the scope you rolled back, and the conditions for trying again completes one cycle.
How to write abort criteria
"Stop if something looks wrong" is not a criterion. The person who has to judge it during the run is already tense, and a vague sentence is no help at that moment. A usable criterion has a number and something to compare it with.
Bad Stop if there are more targets than expected
Usable Stop if the target differs from the 120 in the request
Better Stop if the target is not 120, and print the actual count
Write the three kinds separately. What to check before starting (scope, backup, permissions), what to watch during the run (invariants, error rate, speed), and what to check after it ends (totals, sample comparison). If you mix the three, it becomes unclear what to watch during the run.
Always include at least one invariant. It is a sentence that must be true, such as "The total does not change", "The number of rows with status A becomes 0", or "Not a single row of any other customer changes". Without it, you cannot catch an incident where only the count is right and the content is wrong.
Write down how to stop as well. A criterion whose way of stopping you do not know is not a criterion. For a job that processes rows one at a time, not moving on to the next row is enough, but for a job that changes everything at once, you can stop only if it runs inside a transaction. Building a job in a form that cannot be stopped and then writing abort criteria for it is the most common wasted effort.
And as said before, first ask whether you can split the irreversible change. If you can split it, the weight on the abort criteria is cut in half. It is always better to make the criteria matter less than to write them well.
What you will do in the next lab
You will receive a site where the request said 120 rows but the actual target of the change is 480 rows. You will not stop at writing the abort criteria on paper; you will make the code enforce them.
The grader runs your script directly with three requests and compares the data before and after the run: a request that has to be caught by the scope criterion, a request where an invariant breaks midway and it has to be rolled back, and a request that has to pass.
The last one matters — if it only blocks and gets no work done, that is a failure too.