Sharding Makes It Faster, but the Verdict Can Change
One-line summary
If you split tests into several pieces and run them at the same time, the feedback gets shorter. In exchange, if the split is not deterministic or there are hidden order dependencies between the pieces, the overall verdict changes. Splitting is a trade in which you buy speed and bring in a new failure mode.
Why this is needed
When the feedback slows down, people stop watching the pipeline. Instead of waiting for the result they go off to do other work, and by the time they come back other commits are already stacked on top, so it becomes blurry what broke it. Martin Fowler, borrowing a guideline from extreme programming, presents the 10-minute build as a reasonable target for most projects, and writes that every minute shaved off the build time is a minute saved by every developer on every commit. If the tests take 40 minutes, people end up integrating only once a day, and then the very benefit CI gives disappears.
Splitting is the most direct means of cutting that 40 minutes. Unlike making the server faster or deleting tests, it leaves the tests as they are and reduces only the wall-clock time.
How to split
There are two methods, and both have their place.
Evenly by file. Sort the list of test files and divide it by the number of pieces. The implementation is simple and needs no extra information. In exchange, the time each file takes varies, so the slowest piece decides the total time. If one file takes 10 minutes, no matter how much you increase the pieces, it never goes below 10 minutes.
Balanced by recorded duration. You store how long each test took in the previous run, and divide so that the sums become similar, using those values. The pieces finish at almost the same time, so there is little waste. In exchange, the duration record itself becomes an input, so you must decide where to keep that file and how to update it. For a new test with no record, give a default value and fill it in on the next run.
# 균등 분할: 느린 파일 하나가 전체를 붙잡는다
조각0 [==== ] 2분
조각1 [================== ] 9분 <- 전체 9분
조각2 [===== ] 3분
# 시간 기반 분할: 합이 비슷해진다
조각0 [======== ] 5분
조각1 [======== ] 5분 <- 전체 5분
조각2 [======= ] 4분
The split must be deterministic
Here is the key. For the same commit, the same allocation must come out whenever you run it. If the allocation differs from run to run, three things collapse.
First, you cannot reproduce a failure. A record saying "failed in piece 2" points to a different bundle of tests in the next run. Second, retries lose meaning. Even if you want to rerun just one piece, there is no guarantee that piece contains the same tests. Third, order-dependency problems show up erratically and are mistaken for "tests that fail sometimes".
So the allocation is calculated only from values decided by the commit, not from the run time or random numbers. Sort the list, and read the duration record from a file committed in the repository too. When deciding the pieces by hash, use a stable string such as the test name.
Order dependencies that show up when you split
What splitting most often uncovers is tests that pass when run alone and fail when run together, and the reverse. The cause is usually shared state. A later test leans on a database row, a global variable, a temporary file or a globally registered setting that an earlier test created.
Such tests were broken all along but hid because they happened to always run in the same order. Splitting merely shakes that order and exposes the problem. So it is right to start from the view that a failure right after introducing splitting is usually not a bug of the splitting but a defect that was already there. The way to check is also simple. Run the failing test alone. If it passes alone, there is some state it depends on.
The verdict comes only after merging
The result of each piece is partial information on its own. The verdict comes only after all pieces finish and their results are merged. So the pipeline must have a stage that waits for the pieces, collects the result files and counts them. Each piece uploads its own report as an artifact, and the merge stage downloads and sums them. The GitLab documentation's statement that artifacts are "used to pass intermediate results between stages" describes exactly this use.
What the summation stage must check is not only the pass count. It also checks that no piece is missing and that the total test count is the expected value. If an entire piece died and the rest are all green, a pipeline that sums carelessly judges it green. A piece that ran 0 tests and "succeeded" is the same.
And there is one invariant for the verdict. Even if you change the number of pieces from 3 to 5, the overall verdict must be the same. This is the best way to check whether the split is correct. Run the same commit twice, changing only the number of pieces, and compare the total test count and the failure list.
What it looks like in the field
- If you put the fast tests in the earlier pieces, you learn of a failure within seconds. The whole unit test suite often finishes faster than a single integration test, so just changing the placement order greatly reduces the felt feedback time.
- You increased the number of pieces but the total time is the same. Usually the single slowest file is the wall, or each piece downloads dependencies anew every time so the preparation time exceeds the run time.
- When you hear "only piece 4 keeps failing", first check whether the allocation is deterministic. If the piece number points to different tests each time, that sentence does not even hold.
References
- Continuous integration (fast builds): https://martinfowler.com/articles/continuousIntegration.html
- GitHub Actions matrix: https://docs.github.com/en/actions/using-jobs/using-a-matrix-for-your-jobs
- GitLab job artifacts: https://docs.gitlab.com/ci/jobs/job_artifacts/
What you will do in the next lab
You build a test runner and a splitter yourself with shell and python3. First you split evenly by file and measure how far apart the per-piece times are, and then you read the duration record and split again so that the sums become similar. You split twice from the same commit and compare whether the allocations are the same down to a single character, and compare with a splitter that deliberately mixes in random numbers to see what collapses. Next you put in tests that lean on shared state to reproduce an order dependency, and sum the per-piece reports to check whether the summation stage catches it when one piece is missing. The last step checks whether the overall verdict is the same even when the number of pieces changes.