It worked when I tried again
One-line summary
An intermittent failure is not "something that happens sometimes" but a probability you have not yet measured, and the moment you measure it, how many runs you need before you can say you fixed it is decided along with it.
Why this is needed
"I ran it again and it worked." When this sentence comes out, most investigations stop right there. Since it cannot be reproduced, it cannot be fixed, and since it cannot be fixed, no record is left. The next week when the same thing happens, you start over from the beginning.
But this sentence already contains one piece of data. It failed once and succeeded once. That means one of two trials failed. Of course you cannot call that "a 50% failure rate." That is because there are only two trials. So what to do next is settled — run it more.
There is a separate reason intermittent failures are especially dangerous. They make you judge whether you fixed it from a single success. If you run a problem with a 10% failure rate once without fixing it and it passes, it looks as if you fixed it though you did nothing. The probability of that is 90%.
How it works
1. Write down the number of trials. A failure rate is always a fraction. "It fails sometimes" has neither a numerator nor a denominator. If it failed 24 times out of 200 runs, it is 0.12.
2. Attach an interval to the estimate. 24 out of 200 and 2 out of 20 are both 0.1, but the degree to which you can trust them differs. There are several ways to get an interval for a proportion, and the Wilson score interval is commonly used because it does not collapse even with zero failures or small samples. What matters is that with zero failures, the width of the interval does not become 0 — having seen it zero times does not make the probability 0.
3. Leave the evidence of the failing run on the spot. An intermittent failure cannot be called again. If you write the exit code, standard error, elapsed time, and the gap from the previous run to a file for every trial, the conditions that are true only in the failing runs surface from the table by themselves. Without this record, there is nobody to ask "what did it say back then?"
4. Calculate how many runs you need. The probability that a problem with failure rate p, left unfixed, passes all n runs is (1-p) to the power n. To push that probability down to 1-c or below, you need the following.
(1 - p)^n <= 1 - c 양변에 로그를 취하면
n >= log(1 - c) / log(1 - p)
p=0.25, c=0.95 -> n >= 10.4 -> 11번
p=0.10, c=0.95 -> n >= 28.4 -> 29번
p=0.02, c=0.95 -> n >= 148.3 -> 149번
What this table says is clear. The rarer the problem, the more trials the claim of having fixed it needs. Passing a 2% problem once and saying you fixed it is the same as saying nothing.
What you see in the field
First, observation changes the symptom. If you add logging, attach a debugger, or step through one stage at a time, the symptom disappears. For a problem that depends on timing, the condition does not hold the moment it slows down. This is not that the problem went away but that the method of observation changed the conditions, and the fact that it disappeared is itself a clue. If it disappears when you run it slowly, the cause lies in speed or order.
Second, covering it with retries. Retrying is not the wrong response. But to decide the number of retries, you need the failure rate. If you allow three retries at a 25% failure rate, the probability that all four attempts fail is 0.4%. If you write "three, for now" without this calculation, a number with no basis goes into the operations document.
Third, not quarantining. If a single intermittently failing test blocks the entire deployment pipeline once in every ten runs, people end up learning the habit of ignoring failures. That habit applies equally to real failures. So until it is fixed, you quarantine it, and leave the quarantined list and the date you quarantined it.
Fourth, the basis for having fixed it is a single success. This is why the calculation above exists. Decide first how many times in a row it must pass after the fix, and attach the record of running that many times to the report.
Fifth, not writing down the measurement conditions. Even for the same job, if the failure rate differs between running back to back and running with gaps between, that itself is data pointing to the cause. But if the report says only "failure rate 25%," the next person sees 0% under their own conditions and closes it as "not reproducible." Only if you write how you ran it along with the number of trials can that number be reused.
Reduced to one line: dealing with an intermittent failure is, before it is finding a bug, making a fraction. Once you have a numerator and a denominator, the decisions that follow — how many retries to allow, whether to quarantine, whether you may say it is fixed — all become calculations. Without the fraction, those decisions are all matters of taste.
What really matters in practice
- Turn "sometimes" into a fraction. A failure rate without a denominator is not data.
- Zero is not 0%. Write the interval along with it.
- Evidence can be gathered only on the spot where it failed. Record every trial.
- A claim of having fixed it must be accompanied by a calculated number of trials.
What you will do in the next lab
You receive a job that fails now and then, measure the failure rate and the interval from the number of trials and the number of failures, and extract from the trial records the conditions that are true only in the failing runs. You measure again with gaps to see with your own eyes that observation erases the symptom, and set the retry budget and the quarantine criterion by calculation. Finally, you run the fixed version the calculated number of times and claim "fixed" with evidence.