Not Every Red Build Is the Same Red
One-line summary
Pipeline failures divide into three branches: infrastructure problems, unstable tests and real defects. If you do not separate the branches, people respond to everything with "run it again", and real defects hide inside it.
Why this is needed
If failures happen once every few days, people read the log. If ten times a day, they do not. The retry button gets pressed first, and once it passes it is forgotten. In a team with this habit, even a real defect gets retried two or three times and, if it happens to pass, is merged as it is. This is the moment a red light stops being a signal.
So what is needed is not better logs but classification. Decide first which branch a failure belongs to, and respond differently for each branch.
- Infrastructure. Download failures, name resolution failures, out of disk, out of memory, runner reclamation. Unrelated to the code. The response is retries with an upper bound and capacity adjustment, and you count the occurrences and watch the trend.
- Unstable (flaky). The same commit gives different results. These are tests that lean on time, random numbers, execution order, concurrency or external dependencies. The response is quarantine and repair, and if you cover it with retries, it stays forever.
- Real defect. It fails every time for the same commit. The response is only to fix it or revert it.
The way to tell them apart is surprisingly simple. Run the same commit again: do the results diverge? If they diverge, it is unstable or infrastructure, and if always the same, it is a real defect. So the pipeline must always leave "which commit this run was" in its results.
Turn it into a reproducible failure
To separate unstable from real defect, you must be able to recreate the same conditions. So for every run, you write down the following in the result.
- Random seed. If you shuffle test order or use random data, print the seed in the log and make it possible to run again by feeding in that seed. Randomization that does not leave the seed produces failures that cannot be reproduced.
- Time. Tests that lean on dates break only at month end, at midnight, or in leap years. If you make the time a test uses pinnable, you can revive that condition as it was.
- Environment. Tool versions, locale, time zone, parallelism. When the same commit gives different results, usually one of these differed.
Just leaving these three greatly reduces reports of "I can't reproduce it".
Find when it broke
When looking for the culprit commit of a failure, reading the log backwards takes time proportional to the number of commits. Binary search is logarithmic times a constant. The git bisect documentation says it narrows 675 revisions in roughly 10 steps and 337 in roughly 9 steps. That means narrowing 1000 commits in ten steps.
To automate it there is one condition. The judging script must be deterministic. If it does not give the same answer on the same commit, binary search points at an unrelated commit as the culprit. So you must not run binary search with an unstable test.
git bisect run judges by the script's exit code. The convention is precisely fixed.
0 이 커밋은 정상(good/old)
1..127 이 커밋은 문제 있음(bad/new) 단, 125 는 제외
125 판정할 수 없음 — 이 커밋은 건너뛴다(skip)
128..255 이분 탐색 자체를 중단한다
The reason 125 exists separately is important. If you report cannot-judge, such as a commit that does not build at all, as "has a problem", the culprit gets wrongly driven that way. And 126 and 127 are values the POSIX shell uses for "cannot execute" and "command not found", so 125 was chosen as the largest value available for this purpose. When writing a judging script, be careful not to use something like exit -1. That value becomes 255 and aborts the search altogether.
Feedback time budget
If you do not put a numeric upper bound on feedback time, the pipeline quietly gets slow. Time that grows one step at a time goes unnoticed by everyone, and one day it becomes "it has always taken a while".
It is better to decide in advance what to do when the bound is exceeded. For example, split stages into those that block merging and those that do not and push the slow ones back, split the tests, or skip tests for areas that rarely break depending on the scope of the change. Either way you must record what you gave up. If you do not write down what you gave up, then when an incident happens in that area a few months later, nobody can explain why it was not caught.
What it looks like in the field
- In the first week of counting failure branches, the result is usually that "half is infrastructure". Then it becomes clear that what to fix is not the tests but capacity and retry policy.
- If you keep a list of unstable tests, you start by looking at whether a new failure is on that list, and investigation time drops a lot. If something that was not on the list starts acting unstable, that itself is an event.
- "It worked until yesterday" is usually not yesterday. If you run a binary search, a commit from two weeks ago often comes out. It is only that nobody touched that path in between.
- If you embed which commit an artifact came from, using a value like
git describe, then tracing back "which commit is running right now" during an incident ends in minutes.
References
- git bisect: https://git-scm.com/docs/git-bisect
- git describe: https://git-scm.com/docs/git-describe
- Continuous integration: https://martinfowler.com/articles/continuousIntegration.html
What you will do in the next lab
You create a git repository, stack commits, make it break somewhere in the middle, and then find the culprit with binary search. First you run git bisect start / good / bad by hand and count in how many steps it narrows, and automate the same job with a judging script. You check that the script keeps to the exit code convention, deliberately insert a commit that does not build, and see for yourself how the culprit is wrongly pointed at if you do not skip with 125. Then you put in a judging script whose results diverge and confirm that binary search collapses, and finally write a classification script that counts failure logs by branch and reports.