The one that passes on retry — how many runs prove it is fixed
Goal
You measure the occurrence rate of an intermittent failure by trials and attach an interval to the estimate. You extract from the trial records the conditions that are true only in the failing runs, and measure again with gaps to confirm that observation erases the symptom. Finally, you calculate how many runs you need before you can say you fixed it and prove it by actually running that many.
Why it matters
"I ran it again and it worked" sounds like a sentence that ends the investigation, but it is actually two pieces of data. It failed once and succeeded once. What to do here is run it more to make a fraction. Intermittent failures are dangerous because they make you judge whether you fixed it from a single success. The probability that a problem with a 10% failure rate, left unfixed, passes a single run is 90%. That is why you must first calculate "how many times in a row must it pass." What is hard in this lab is not the loop but the basis of the judgment. If you write only the proportion without the number of trials, you cannot tell 2 out of 20 from 20 out of 200, and if you do not leave the evidence of the failing run on the spot, there is nobody to ask later. The grader is not shaken by randomness. It runs your tool against a deterministic target the grader built (a command that fails exactly once in every five runs) and compares the numbers, and the records and calculations you leave are recomputed by the grader with the same formulas and matched.
Steps
- Create and run /root/flaky/gen_flaky.py to create /root/flaky/job.py.
- Use /root/flaky/measure.py to run job.py at least 200 times and write the failure rate and the Wilson interval in /root/flaky/rate.json.
- Use /root/flaky/plan.py and /root/flaky/runs_needed.json to work out the number of consecutive passes needed for each confidence level.
- Use /root/flaky/collect.py to leave evidence for every trial and create /root/flaky/facts.jsonl, and write what is true only when it fails in /root/flaky/only_when_fail.json.
- Measure again with gaps between trials and write the two measurements side by side in /root/flaky/observer.json.
- In /root/flaky/policy.json, set the retry budget, the quarantine criterion, and the number of consecutive passes needed for a green verdict, by calculation.
- Run the fixed version the calculated number of times and leave the evidence in /root/flaky/proof.json.
- Report in /root/flaky/summary.json and /root/flaky/flaky_report.md in four sections.
Notes
- Job contract:
python3 /root/flaky/job.py [--fixed] [--state <디렉터리>] [--seed <정수>](where the placeholders are a directory and an integer) outputs one line of JSON and exit code 0 on success, and one line of standard error and exit code 3 on failure. - Measurement contract:
python3 /root/flaky/measure.py --cmd "<명령>" --trials <횟수> [--gap-ms <밀리초>] --out <json>(with the command, count, milliseconds, and output path filled in) outputs one JSON object containing command, trials, failures, rate, ci_low, ci_high, and gap_ms.--gap-msis the rest time before each trial — it rests before the first trial too, because you do not know what ran just before. - Wilson score interval (95%, z=1.96): the center is
(p + z²/2n) / (1 + z²/n)and the half-width isz/(1 + z²/n) × √(p(1-p)/n + z²/4n²). Even with zero failures, the width does not become 0. - Planning contract:
python3 /root/flaky/plan.py --rate <실패율> --confidence <신뢰수준>(with the failure rate and the confidence level filled in) outputs JSON containing rate, confidence, and runs.runsis the smallest integer satisfying(1-p)^n <= 1-c, that is,log(1-c)/log(1-p)rounded up. - Collection contract:
python3 /root/flaky/collect.py --cmd "<명령>" --trials <횟수> [--gap-ms <밀리초>] --out <jsonl>(with the command, count, milliseconds, and output path filled in) writes trial, exit_code, outcome, elapsed_ms, stderr, and stdout, one trial per line. - Retry budget: the probability that every attempt fails as you try one more time is a power of p. Use the smallest r satisfying
p^(r+1) <= 0.01as the budget. - Common mistakes: not writing the number of trials, writing zero failures as a failure rate of 0, reporting that it disappeared when measured with observation attached, and taking a single success as the basis of having fixed it.
- Assumption of this lab: the 95% confidence level and the 1% retry budget criterion are values set in this lab. In the field they vary with the cost of failure.
- Do not build a load test. The budget for one grading is 60 seconds and the Pod has 2 cores. For the slow measurement in step 5, 20 runs are enough (it takes about a minute because of the gaps).
Get the intermittently failing job in hand
Create and run /root/flaky/gen_flaky.py to create /root/flaky/job.py. Then run job.py by hand several times and see with your own eyes that it sometimes works and sometimes does not.
Just save this script as it is and run it. Run the created job.py about ten times in a row. Runs with exit code 0 and runs with exit code 3 come out mixed. Do not count yet how often it is — that is step 2.
Turn sometimes into a fraction
Create /root/flaky/measure.py so that it runs job.py at least 200 times, and write command, trials, failures, rate, ci_low, ci_high, and gap_ms in /root/flaky/rate.json. The interval is a 95% Wilson score interval.
Count a failure as a non-zero exit code. If you write only the proportion, you cannot tell 2 out of 20 from 20 out of 200, so leave the number of trials along with it. The two formulas for the Wilson interval are in the notes — the reason to use this interval is that the width does not become 0 even with zero failures.
How many runs before you can say you fixed it
Create /root/flaky/plan.py, and in /root/flaky/runs_needed.json write observed_rate, a table containing runs for each of the confidence levels 0.9, 0.95, and 0.99, and chosen (0.95).
The probability that it passes all n runs without being fixed is (1-p) to the power n. The smallest n that pushes that probability down to 1-c or below is the answer, and if you take logarithms it comes out in one line. Do not forget to round up — you cannot run a fractional number of times.
Leave evidence on the spot where it failed
Run /root/flaky/collect.py at least 60 times to create /root/flaky/facts.jsonl, and in /root/flaky/only_when_fail.json write trials, failures, exit_codes_on_failure, exit_codes_on_success, token, and failure_trials. token is a word that is in the standard error of every failure and in no success.
An intermittent failure cannot be called again. You must write the exit code and standard error on that trial's line, so that later there is someone to ask. The token is visible right away if you skim the failure records by eye — it is a single word in capital letters. It becomes a token only if that word is absent from the success records.
Observation erases the symptom
Measure again at least 20 times with a gap of at least 2000 milliseconds between trials, and write the two measurements fast and slow (trials, failures, rate, gap_ms) and changed in /root/flaky/observer.json. When measured slowly, there must be 0 failures.
It is common for the symptom to disappear when you add logging and step through one stage at a time. That is because a timing-dependent condition does not hold the moment it slows down. The fact that it disappeared is itself a clue — it means the cause lies in speed or order. You can make the same thing with --gap-ms of measure.py.
Calculate the retry budget and the quarantine criterion
In /root/flaky/policy.json, write observed_rate, confidence (0.95), green_runs_required, retry_budget, and quarantine_after. retry_budget is the smallest r satisfying p^(r+1) <= 0.01, and quarantine_after is an integer of at least 1.
Retrying is not the wrong response, but the count must have a basis. If you know the failure rate, you can calculate 'the probability that all fail,' and the budget comes from that. If you do not quarantine, people learn the habit of ignoring failures, and that habit applies to real failures too.
Claim it is fixed, with evidence
Run the fixed version (--fixed) at least the number of times calculated in step 3, and write command, required, trials, and failures in /root/flaky/proof.json. failures must be 0 and trials must be at least required.
A single pass is not evidence. That is because the probability that a problem with failure rate p passes without being fixed is already 1-p. Take chosen.runs from step 3 as it is, run that many times, and leave that record in a file. The grader runs the same command again itself.
Report the probability on one page
In /root/flaky/summary.json, write trials, failures, rate, ci_low, ci_high, token, green_runs_required, proof_trials, and observer_changed, and in /root/flaky/flaky_report.md, report in four sections: ## 얼마나 자주인가 ## 실패할 때만 참인 것 ## 관측이 바꾼 것 ## 고쳤다고 말하려면 (in order, these mean: how often, what is true only when it fails, what observation changed, and what it takes to say it is fixed).
The value of the report lies not in the conclusion but in the numbers. Write the number of trials, the number of failures, the interval, and the number of consecutive passes needed. The fact that it disappeared when observation was attached is also a clue, so write it along with them.