TT Lab
Get started
Learn Learning paths Courses

GitLab CI/CD

What the Runner Actually Does

Continue in TT Lab

In one sentence

GitLab only reads the configuration and builds the job list, and what actually runs those jobs is a separate process, the runner. So without a runner, a pipeline stays in a waiting state forever, without an error.

Why this was needed

The block that people handling pipelines for the first time run into most often is not a syntax error but the situation "the job was created but does not start". There is no red error on the screen, and the job only says pending. This state does not mean the configuration is wrong but that there is no runner to pick up that job, and if you don't know the structure in which the server and the executor are separated, you keep staring at the configuration file.

The separation itself is an intended design. Builds are heavy and dangerous. They run arbitrary code, fill the disk, and go out over the network. If you do such things in the same process as the server, the server dies with it. If you move execution outside, you can add as many runners as needed, give each team a different environment, and GitLab stays fine even if one runner breaks.

How it works

After a runner is registered with GitLab, it periodically asks "is there a job I can take". It is not that the server pushes; the runner pulls, so even if the runner is behind a firewall, there is no need to open an inbound port.

What it takes is decided by tags. If you write tags on a job, only runners with that tag pick up that job. If you don't write a tag on a job that needs one, or if there is not a single runner with that tag, the job keeps waiting. This is the typical cause of the pending mentioned above.

How it runs after picking up is decided by the executor. shell runs as it is on the machine where the runner is installed, docker starts a new container for each job, and kubernetes creates a Pod for each job. The latter two give a clean environment each time, so reproducibility is good, but in exchange nothing remains between jobs. This is exactly why the artifacts and cache you learned earlier are needed — they are devices for handing over only what is needed while keeping the isolation.

During a run, GitLab puts in a lot of variables that start with CI_. Things like CI_COMMIT_BRANCH, CI_COMMIT_SHA, CI_PIPELINE_ID, and CI_JOB_ID. The conditions of rules are written with these variables, and artifact names and image tags are also made with them. The habit of putting the commit SHA in the image tag comes from here.

What you see in the field

The number most often argued over in runner operations is the concurrency. If it is too low, jobs queue up, and if it is too high, the machine starts swapping and all jobs slow down together. And a slowed pipeline calls for retries, and retries create load again.

Another is the boundary between shared runners and dedicated runners. If you run a deployment job on a shared runner that anyone can use, it becomes possible for other projects using that runner to see the place where the deployment credentials passed. So it is the basic practice to separate deployment jobs onto runners dedicated to protected branches.

In addition, this lab environment has neither a GitLab server nor a runner. The first two labs handled how jobs are created and in what order they queue up, using the configuration and an interpreter you wrote yourself, and the labs after that run with gitlab-ci-local, which interprets the configuration by the same rules as GitLab and actually runs the jobs in a shell. Behaviors that require a server, such as runner registration, tag matching, and protected variables, are not reproduced, and each lab states that limit.

What to look at next

We look at the places where pipelines cause incidents. We cover the paths through which secret values leak, the risk of pulling in someone else's configuration, and the minimum devices that make a deployment revertible.