Stages Queue Up; needs Builds a DAG
In one sentence
stages lines jobs up into slots so that the next slot starts only after the previous slot has completely finished, and needs ignores those slot walls and specifies only "what I actually have to wait for", turning the pipeline into a DAG.
Why this was needed
The stage model has the big advantage of being easy to understand. When build finishes completely, test starts, and when test finishes completely, deploy starts. The problem is that this simplicity turns into waste.
Suppose the test stage has five jobs and one of them is a 12-minute integration test. The other four finish in a minute, but the jobs in the deploy stage wait the full 12 minutes. In reality, the deployment may need only the integration test result, but the stage model has no way to express that fact. The real dependencies between jobs are much sparser than slots, and slots round that sparse relationship up into a dense one.
The opposite side is even more frustrating. A lint job that only needs to look at the source cannot even start until the build finishes, just because it comes after the build stage. The developer waits for the build time just to learn about one typo. When the total pipeline time exceeds 15 minutes, people stop waiting for the results and go off to do other things, and the feedback loop collapses, and this kind of pointless waiting is a large share of those 15 minutes.
How it works
If you write needs on a job, that job ignores the stage order and waits only for the listed jobs. Three rules work together here.
First, a job with no needs at all stays as before. It starts only when every job in the stage before its own has finished. So the stage style and the DAG style may be mixed in one file.
Second, needs: [] means "wait for nothing". A missing key and an empty list have exactly opposite meanings. This single line pulls a lint job to the very front of the pipeline.
Third, needs is not limited to pointing at earlier stages. It can also point at a job in the same stage, and then an order arises even within the same slot. Conversely, if it points at a job in a later stage, it becomes a cycle and GitLab rejects the configuration.
Combine these three rules and the pipeline becomes a directed acyclic graph, and the execution order is decided by a topological sort of the graph. If you want to see the execution plan with your eyes, count it in waves. Jobs with nothing to wait for start together in the first wave, and the jobs that can start only once those finish become the second wave. The number of waves is the minimum depth of that pipeline, and it is the lower bound of time that does not shrink no matter how many runners you add.
There is one practical constraint. There is an upper limit on the number of needs a job can have (50 by default), so attempts to turn a pipeline of hundreds of jobs into a complete DAG usually stop halfway. In that case, it is better to pick only the few that are bottlenecks, apply needs to them, and leave the rest to stages.
What you see in the field
The first surprise for a team that has switched to a DAG is not that time shrinks but that the dependencies get documented. To write needs, you have to answer what this job is waiting for, and when you actually try to write it, several waits come out for which nobody knows the reason. Most of those waits could be deleted.
There are incidents in the opposite direction too. A job moved forward with needs was using an earlier job's output, but because that dependency was not written in needs, it fails saying the file is missing. This failure does not reproduce when the runners are idle and appears only when they are busy, so it is easily mistaken for a flaky test.
What to look at next
Even in the same pipeline, different jobs must be created depending on the branch and on whether it is a merge request. We look at rules, which handles that choice, in the next article.