TT Lab
Get started
Learn Learning paths Courses

CI/CD Pipelines

Ten red builds and nobody read a single log

Continue in TT Lab

Goal

You build an investigation flow from start to finish: dividing pipeline failures into branches with a table, sending each branch to a different response, proving instability by measurement, recreating a failed run from the record alone, and automatically finding the culprit commit with a deterministic judging script.

Why it matters

If failures happen ten times a day, people do not read the logs. The retry button gets pressed first, and once it passes it is forgotten. In a pipeline with that habit, even a real defect is retried two or three times and, if it happens to pass, merged as it is, and the red light stops being a signal. So what is needed is not better logs but branches. An infrastructure problem may be retried, an unstable test stays forever if covered with retries, and a real defect gives the same answer however many times you run it, so retrying only wastes time. Only after you have decided the branch does the question 'when did it break?' hold, and the binary search that answers it can be trusted only when the judgment is deterministic. This is because a single wrong judgment drives the entire remaining search into the wrong range. This lab builds that whole chain — dividing with a table, proving with measurement, reviving from a record, and finding the culprit with a judging script that speaks only through exit codes.

Steps

  1. In /root/triage/rules.tsv, write the classification rules as a table. It has three tab-separated columns, 규칙id<TAB>갈래<TAB>확장정규식 (rule id, category, extended regular expression), and the first match from the top wins. Put six rules with these ids and categories — dns (infra), disk (infra), oom (infra), timeout (flaky), assert (defect), syntax (defect). Then, in /root/triage/logs/, write six sample failure logs yourself, from run-01.log to run-06.log. In order, their contents must show a name resolution failure (dns), out of disk (disk), dying from out of memory (oom), a timeout (timeout), an assertion failure (assert) and a syntax error (syntax). Finally, create /root/triage/classify.sh <로그파일> (log file). The location of the rules table can be changed with the environment variable RULES_FILE, and the default is /root/triage/rules.tsv. It prints on one line the category as the first word and the rule id as the second word. If no rule matches, it prints unknown -. If the log file cannot be read, it prints nothing to standard output and ends with exit code 2.
  2. In /root/triage/policy.tsv, write the response by category as a table. It has three tab-separated columns, 갈래<TAB>권장대응<TAB>재시도가능 (category, recommended response, retry allowed) and four lines — infra requeue yes, flaky measure no, defect bisect no, unknown read no (tabs between the columns). Then fix /root/triage/classify.sh so that it prints four words on one line, <갈래> <규칙id> <권장대응> <재시도가능> (category, rule id, recommended response, retry allowed). The location of the policy table can be changed with the environment variable POLICY_FILE, and the default is /root/triage/policy.tsv. And make it also say the category through the exit code — infra is 0, flaky is 3, defect is 4, unknown is 5, and if the log cannot be read, 2. Finally, run the six logs in /root/triage/logs/ in turn and save their output as it is, as six lines, to /root/triage/triage.txt (in the order run-01 to run-06).
  3. In /root/triage/tests/, create three test scripts — always-pass.sh (always 0), always-fail.sh (always non-zero) and flip.sh (fails only on odd-numbered runs). Put the value that flip.sh counts next to its own file (under $(dirname "$0")) so that it follows along even when the script is copied and moved. Then create /root/triage/flaky-probe.sh <시험스크립트> <횟수> (test script, count). It runs the script it receives that many times and prints one line runs=<횟수> pass=<성공> fail=<실패> verdict=<판정> (count, passes, failures, verdict), and the verdict is one of stable-pass, flaky and stable-fail. The exit code is 0 for stable-pass, 3 for flaky and 4 for stable-fail, and 2 if there is no argument or the count is not an integer of 1 or more. Finally, save the output of running each of the three scripts 8 times to /root/triage/flaky-report.txt as three lines (in the order always-pass, flip, always-fail).
  4. Create /root/triage/tests/seed-test.sh. It is a test that reads the environment variable SEED and gives the same result every time for the same seed but with success and failure diverging depending on the seed. If SEED is empty, it ends with exit code 2. Next create /root/triage/record-failure.sh <시험스크립트> <기록파일> (test script, record file). It puts a different value each time into the environment variable SEED and runs up to 30 times, stops at the first failure, writes five lines in the record file as SEED=, TZ=, LC_ALL=, SCRIPT_SHA256= (the first 64 digits of the test script's sha256) and EXIT_CODE=, and ends with 0. If there is no failure within 30 runs, it leaves no record and ends with a non-zero value. Finally create /root/triage/replay.sh <기록파일> <시험스크립트> (record file, test script). If the SCRIPT_SHA256 in the record differs from the hash of the current script, it does not run the test and ends with exit code 2, and if the same, it puts the record's SEED, TZ and LC_ALL and runs the test, then ends with that test's exit code. Chain the three and create /root/triage/repro/case.env — it is record-failure.sh /root/triage/tests/seed-test.sh /root/triage/repro/case.env.
  5. Create a practice git repository at /root/triage/repo. There are 20 commits and the subjects are change 1 through change 20. Each commit contains app/rate.py and app/notes.txt, and python3 app/rate.py outputs 100 from commit 1 to commit 11 and 250 from commit 12 (commit 12 is the commit that introduced the defect). The Pod has no git identity, so first set git config user.name and git config user.email inside the repository. In /root/triage/expected.txt, write the single line with the expected value 100. Then create /root/triage/oracle.sh. It runs at the root of the checked-out working tree and ends with 0 if the output of python3 app/rate.py equals the expected value and 1 if not. The location of the expected value file can be changed with the environment variable EXPECT_FILE, and the default is /root/triage/expected.txt. It prints nothing to standard output.
  6. In /root/triage/repo, mark the first commit as good and HEAD as bad, and find the culprit commit with git bisect run. Use /root/triage/oracle.sh for the judging. After finding it, write two lines in /root/triage/bisect/culprit.txt — culprit=<40자리 커밋 해시> (40-character commit hash) and subject=<그 커밋의 제목> (that commit's subject). And write three lines in /root/triage/bisect/steps.txt — revisions=<good 뒤부터 bad 까지의 커밋 수> (the number of commits from after good up to bad), tests=<판정 스크립트가 실제로 불린 횟수> (the number of times the judging script was actually called) and reason=<왜 그 횟수인지 한 줄 설명> (a one-line explanation of why that count) (20 characters or more). When finished, return the repository to its original branch with git bisect reset.
  7. Create a second repository at /root/triage/repo-skip. It has 20 commits with subjects of the same shape, but this time commit 10's app/rate.py has a syntax error so it cannot even run, and from commit 18 it outputs 250 (from commit 1 to commit 17, all are 100 except commit 10). Then create /root/triage/oracle-skip.sh. It is like oracle.sh, but if app/rate.py is missing or the syntax does not pass, it ends with exit code 125. In the same repository, mark the first commit as good and HEAD as bad and run the binary search twice — once with /root/triage/oracle.sh and once with /root/triage/oracle-skip.sh. Write the result in four lines in /root/triage/bisect/skip-report.txt — naive=<oracle.sh 가 지목한 40자리 해시> (the 40-character hash that oracle.sh pointed at), skip=<oracle-skip.sh 가 지목한 40자리 해시> (the 40-character hash that oracle-skip.sh pointed at), true=<진짜 범인의 40자리 해시> (the 40-character hash of the true culprit) and limit=<건너뛰기의 한계를 적은 한 줄> (one line stating the limit of skipping) (20 characters or more). Do not forget git bisect reset when you finish.
  8. Create /root/triage/investigate.sh <로그파일> <저장소> <보고서파일> (log file, repository, report file). It classifies the log with /root/triage/classify.sh, and only when the category is defect, it clones the repository with git clone and binary-searches from the first commit to HEAD with /root/triage/oracle-skip.sh to find the culprit. The repository it receives must not change by a single character, and the files that someone was fixing in that repository must also stay as they are. The report is five lines — category=, action=, culprit= (- if it is not defect or it could not be found), subject= (likewise) and reproduce= (one line on how to cause it again). If the parent directory of the report file does not exist, it creates it, and it ends with 0 on normal completion and with 2 if arguments are missing or the log or repository cannot be read. After creating it, run it twice — with /root/triage/logs/run-05.log and /root/triage/repo-skip, create /root/triage/report/case-defect.txt, and with /root/triage/logs/run-01.log and /root/triage/repo-skip, create /root/triage/report/case-infra.txt.

Notes

There are ten red lights and nobody has read them

In /root/triage/rules.tsv, write the classification rules as a table. It has three tab-separated columns, 규칙id<TAB>갈래<TAB>확장정규식 (rule id, category, extended regular expression), and the first match from the top wins. Put six rules with these ids and categories — dns (infra), disk (infra), oom (infra), timeout (flaky), assert (defect), syntax (defect). Then, in /root/triage/logs/, write six sample failure logs yourself, from run-01.log to run-06.log. In order, their contents must show a name resolution failure (dns), out of disk (disk), dying from out of memory (oom), a timeout (timeout), an assertion failure (assert) and a syntax error (syntax). Finally, create /root/triage/classify.sh <로그파일> (log file). The location of the rules table can be changed with the environment variable RULES_FILE, and the default is /root/triage/rules.tsv. It prints on one line the category as the first word and the rule id as the second word. If no rule matches, it prints unknown -. If the log file cannot be read, it prints nothing to standard output and ends with exit code 2.

The reason to put the rules in a table instead of code is that when a new failure shape appears, the place to fix must be data rather than the script for reviews and history to remain. Read tab-separated lines with while IFS=$'\t' read -r a b c. If the last line has no newline, read drops that line, so it is safer to add || [ -n "$a" ]. Ask about extended regular expressions with grep -Eq -- "$pattern". The grader also runs this script with a rules table it made itself and logs it made itself — if you answer by looking at the log file name, you get caught then.

Failures you may retry and failures you must never retry

In /root/triage/policy.tsv, write the response by category as a table. It has three tab-separated columns, 갈래<TAB>권장대응<TAB>재시도가능 (category, recommended response, retry allowed) and four lines — infra requeue yes, flaky measure no, defect bisect no, unknown read no (tabs between the columns). Then fix /root/triage/classify.sh so that it prints four words on one line, <갈래> <규칙id> <권장대응> <재시도가능> (category, rule id, recommended response, retry allowed). The location of the policy table can be changed with the environment variable POLICY_FILE, and the default is /root/triage/policy.tsv. And make it also say the category through the exit code — infra is 0, flaky is 3, defect is 4, unknown is 5, and if the log cannot be read, 2. Finally, run the six logs in /root/triage/logs/ in turn and save their output as it is, as six lines, to /root/triage/triage.txt (in the order run-01 to run-06).

The reason for also giving the category as an exit code is so that later automation does not have to parse the string again. The last step of this lab uses that value. Only infrastructure is worth retrying — unstable stays forever if covered with retries, and a real defect gives the same answer however many times you run it, so it only wastes time. The grader runs it by passing a policy table it made itself through POLICY_FILE. If you hard-code the response words inside the script, you get caught then.

You cannot call something unstable just because it failed once

In /root/triage/tests/, create three test scripts — always-pass.sh (always 0), always-fail.sh (always non-zero) and flip.sh (fails only on odd-numbered runs). Put the value that flip.sh counts next to its own file (under $(dirname "$0")) so that it follows along even when the script is copied and moved. Then create /root/triage/flaky-probe.sh <시험스크립트> <횟수> (test script, count). It runs the script it receives that many times and prints one line runs=<횟수> pass=<성공> fail=<실패> verdict=<판정> (count, passes, failures, verdict), and the verdict is one of stable-pass, flaky and stable-fail. The exit code is 0 for stable-pass, 3 for flaky and 4 for stable-fail, and 2 if there is no argument or the count is not an integer of 1 or more. Finally, save the output of running each of the three scripts 8 times to /root/triage/flaky-report.txt as three lines (in the order always-pass, flip, always-fail).

Whether the results diverge under the same conditions — this is both the definition and the way to tell instability. A single failure does not tell you which branch it belongs to. If you put the state of flip.sh in /tmp or a pinned path, then when you make a copy it shares state with the original and the measurement goes off. For the count check, case "$N" in ''|*[!0-9]*) is short. The grader copies the student's tests/ wholesale and runs it in the copy.

A failure you cannot reproduce cannot be investigated

Create /root/triage/tests/seed-test.sh. It is a test that reads the environment variable SEED and gives the same result every time for the same seed but with success and failure diverging depending on the seed. If SEED is empty, it ends with exit code 2. Next create /root/triage/record-failure.sh <시험스크립트> <기록파일> (test script, record file). It puts a different value each time into the environment variable SEED and runs up to 30 times, stops at the first failure, writes five lines in the record file as SEED=, TZ=, LC_ALL=, SCRIPT_SHA256= (the first 64 digits of the test script's sha256) and EXIT_CODE=, and ends with 0. If there is no failure within 30 runs, it leaves no record and ends with a non-zero value. Finally create /root/triage/replay.sh <기록파일> <시험스크립트> (record file, test script). If the SCRIPT_SHA256 in the record differs from the hash of the current script, it does not run the test and ends with exit code 2, and if the same, it puts the record's SEED, TZ and LC_ALL and runs the test, then ends with that test's exit code. Chain the three and create /root/triage/repro/case.env — it is record-failure.sh /root/triage/tests/seed-test.sh /root/triage/repro/case.env.

If you cannot cause the failure again, there is no way to tell whether it is unstable or a real defect. So the record contains only what is needed to revive that run — the seed, time zone, locale, and a hash to confirm that the input is still the same. The reason to compare the hash is that a run after the input has changed is not a reproduction but a new experiment. If you let it pass silently, the wrong conclusion 'reproduced' remains. Make the seed by concatenating $RANDOM or mixing in the loop variable. If you run 30 times with the same seed, you just get the same result 30 times — the grader counts how many different seeds there were.

First build the judging script that binary search can trust

Create a practice git repository at /root/triage/repo. There are 20 commits and the subjects are change 1 through change 20. Each commit contains app/rate.py and app/notes.txt, and python3 app/rate.py outputs 100 from commit 1 to commit 11 and 250 from commit 12 (commit 12 is the commit that introduced the defect). The Pod has no git identity, so first set git config user.name and git config user.email inside the repository. In /root/triage/expected.txt, write the single line with the expected value 100. Then create /root/triage/oracle.sh. It runs at the root of the checked-out working tree and ends with 0 if the output of python3 app/rate.py equals the expected value and 1 if not. The location of the expected value file can be changed with the environment variable EXPECT_FILE, and the default is /root/triage/expected.txt. It prints nothing to standard output.

If you put the judging criterion inside the repository, the criterion changes along with everything else as you move between commits, and nothing can be judged. So keep the expected value outside the repository and open up its location through an environment variable — that way you can use the same judging script for other repositories too. The reason to keep standard output empty is that binary search reads only the exit code. If you mix a verdict onto the screen, the automating side ends up parsing that string again. Write the loop that makes the history briefly with git add -A and git commit -qm. The grader steps through the commits one by one in a clone and checks that the place where it goes from normal to having a problem is exactly one.

Narrow nineteen commits in four tries

In /root/triage/repo, mark the first commit as good and HEAD as bad, and find the culprit commit with git bisect run. Use /root/triage/oracle.sh for the judging. After finding it, write two lines in /root/triage/bisect/culprit.txt — culprit=<40자리 커밋 해시> (40-character commit hash) and subject=<그 커밋의 제목> (that commit's subject). And write three lines in /root/triage/bisect/steps.txt — revisions=<good 뒤부터 bad 까지의 커밋 수> (the number of commits from after good up to bad), tests=<판정 스크립트가 실제로 불린 횟수> (the number of times the judging script was actually called) and reason=<왜 그 횟수인지 한 줄 설명> (a one-line explanation of why that count) (20 characters or more). When finished, return the repository to its original branch with git bisect reset.

The reason to keep the judging script outside the repository shows up here — binary search moves between commits, so if you put it inside the repository, that script changes from commit to commit too. Count the number of commits in the range with git rev-list --count <good>..<bad>. To get the number of calls, count the lines that git bisect run prints on the screen, running ..., or wrap the judging script in a shell that logs one line per call and count those. Write in reason= why the count does not follow the number of commits but is proportional to its logarithm. The grader runs the same binary search itself in a clone and checks the answer and the count.

One commit that cannot even run changed the culprit

Create a second repository at /root/triage/repo-skip. It has 20 commits with subjects of the same shape, but this time commit 10's app/rate.py has a syntax error so it cannot even run, and from commit 18 it outputs 250 (from commit 1 to commit 17, all are 100 except commit 10). Then create /root/triage/oracle-skip.sh. It is like oracle.sh, but if app/rate.py is missing or the syntax does not pass, it ends with exit code 125. In the same repository, mark the first commit as good and HEAD as bad and run the binary search twice — once with /root/triage/oracle.sh and once with /root/triage/oracle-skip.sh. Write the result in four lines in /root/triage/bisect/skip-report.txt — naive=<oracle.sh 가 지목한 40자리 해시> (the 40-character hash that oracle.sh pointed at), skip=<oracle-skip.sh 가 지목한 40자리 해시> (the 40-character hash that oracle-skip.sh pointed at), true=<진짜 범인의 40자리 해시> (the 40-character hash of the true culprit) and limit=<건너뛰기의 한계를 적은 한 줄> (one line stating the limit of skipping) (20 characters or more). Do not forget git bisect reset when you finish.

125 means 'this commit cannot be judged'. If you fold cannot-judge into 1 (has a problem), the search converges on the area before it and points at an innocent commit as the one with confidence — the side that does not say it does not know is worse. If you use py_compile to check only the syntax, __pycache__ remains in the working tree and blocks the next checkout. python3 -c 'import ast,sys; ast.parse(open(sys.argv[1]).read())' leaves nothing behind. The true culprit is the place where, going one by one from the oldest commit, it first becomes having a problem — if you clone the repository and sweep through it, you do not touch the original. In limit=, write how far the search can answer when the skipped commit is adjacent to the culprit.

Make one log in produce the culprit commit

Create /root/triage/investigate.sh <로그파일> <저장소> <보고서파일> (log file, repository, report file). It classifies the log with /root/triage/classify.sh, and only when the category is defect, it clones the repository with git clone and binary-searches from the first commit to HEAD with /root/triage/oracle-skip.sh to find the culprit. The repository it receives must not change by a single character, and the files that someone was fixing in that repository must also stay as they are. The report is five lines — category=, action=, culprit= (- if it is not defect or it could not be found), subject= (likewise) and reproduce= (one line on how to cause it again). If the parent directory of the report file does not exist, it creates it, and it ends with 0 on normal completion and with 2 if arguments are missing or the log or repository cannot be read. After creating it, run it twice — with /root/triage/logs/run-05.log and /root/triage/repo-skip, create /root/triage/report/case-defect.txt, and with /root/triage/logs/run-01.log and /root/triage/repo-skip, create /root/triage/report/case-infra.txt.

If the investigation script runs binary search directly in someone else's repository, it takes away the place that person was working in. Run it in a temporary directory you git cloned and clean it up with trap ... EXIT. The judging script reads the location of the expected value file from an environment variable — the investigation script should just pass that variable along without clearing it. The grader runs it passing its own expected value file. Find the first commit with git rev-list --max-parents=0 HEAD. In the infrastructure branch, it must not run binary search at all — digging through commits for a failure unrelated to code is a waste of time.