TT Lab
Get started
Learn Learning paths Courses

CI/CD Pipelines

Not Every Red Build Is the Same Red

Continue in TT Lab

One-line summary

Pipeline failures divide into three branches: infrastructure problems, unstable tests and real defects. If you do not separate the branches, people respond to everything with "run it again", and real defects hide inside it.

Why this is needed

If failures happen once every few days, people read the log. If ten times a day, they do not. The retry button gets pressed first, and once it passes it is forgotten. In a team with this habit, even a real defect gets retried two or three times and, if it happens to pass, is merged as it is. This is the moment a red light stops being a signal.

So what is needed is not better logs but classification. Decide first which branch a failure belongs to, and respond differently for each branch.

The way to tell them apart is surprisingly simple. Run the same commit again: do the results diverge? If they diverge, it is unstable or infrastructure, and if always the same, it is a real defect. So the pipeline must always leave "which commit this run was" in its results.

Turn it into a reproducible failure

To separate unstable from real defect, you must be able to recreate the same conditions. So for every run, you write down the following in the result.

Just leaving these three greatly reduces reports of "I can't reproduce it".

Find when it broke

When looking for the culprit commit of a failure, reading the log backwards takes time proportional to the number of commits. Binary search is logarithmic times a constant. The git bisect documentation says it narrows 675 revisions in roughly 10 steps and 337 in roughly 9 steps. That means narrowing 1000 commits in ten steps.

To automate it there is one condition. The judging script must be deterministic. If it does not give the same answer on the same commit, binary search points at an unrelated commit as the culprit. So you must not run binary search with an unstable test.

git bisect run judges by the script's exit code. The convention is precisely fixed.

0          이 커밋은 정상(good/old)
1..127     이 커밋은 문제 있음(bad/new)   단, 125 는 제외
125        판정할 수 없음 — 이 커밋은 건너뛴다(skip)
128..255   이분 탐색 자체를 중단한다

The reason 125 exists separately is important. If you report cannot-judge, such as a commit that does not build at all, as "has a problem", the culprit gets wrongly driven that way. And 126 and 127 are values the POSIX shell uses for "cannot execute" and "command not found", so 125 was chosen as the largest value available for this purpose. When writing a judging script, be careful not to use something like exit -1. That value becomes 255 and aborts the search altogether.

Feedback time budget

If you do not put a numeric upper bound on feedback time, the pipeline quietly gets slow. Time that grows one step at a time goes unnoticed by everyone, and one day it becomes "it has always taken a while".

It is better to decide in advance what to do when the bound is exceeded. For example, split stages into those that block merging and those that do not and push the slow ones back, split the tests, or skip tests for areas that rarely break depending on the scope of the change. Either way you must record what you gave up. If you do not write down what you gave up, then when an incident happens in that area a few months later, nobody can explain why it was not caught.

What it looks like in the field

References

What you will do in the next lab

You create a git repository, stack commits, make it break somewhere in the middle, and then find the culprit with binary search. First you run git bisect start / good / bad by hand and count in how many steps it narrows, and automate the same job with a judging script. You check that the script keeps to the exit code convention, deliberately insert a commit that does not build, and see for yourself how the culprit is wrongly pointed at if you do not skip with 125. Then you put in a judging script whose results diverge and confirm that binary search collapses, and finally write a classification script that counts failure logs by branch and reports.