TT Lab
Get started
Learn Learning paths Courses

GitLab CI/CD

One slow e2e job made docs publishing wait eight seconds

Continue in TT Lab

Goal

You actually run a pipeline that has slow jobs mixed in, measure how stages and needs change when each job starts, and confirm from the run results how artifact passing, waiting for conditional jobs, and allowing failures behave.

Why it matters

The stage approach is easy to understand, but the slowest job holds up the start of all later jobs. If you make jobs wait only for what they need with needs, the pipeline gets faster, but in exchange you must write exactly, job by job, what they wait for and what they receive. If you wait for a job that may drop out through rules, the pipeline is not created at all, and if you allow failures too broadly, a real incident passes by in green. These choices become clear when you look at times and results rather than when you read the documentation.

Steps

  1. Make /root/glci-dag a git repository and in .gitlab-ci.yml put stages [build, test, deploy] and four jobs: compile (build; after sleep 3, it creates the file bin/app and uploads bin/ as artifacts), unit (test; test -f bin/app and then echo unit-ok), e2e (test; after sleep 8, echo e2e-ok), and publish-docs (deploy; test -f bin/app and then echo docs-published). Run it with gitlab-ci-local --shell-isolation --no-artifacts-to-source --timestamps and see that publish-docs starts only after the slow e2e has finished.
  2. To publish-docs, add needs: [compile]. When you run it, publish-docs should start before e2e finishes, and it should still receive compile's output (bin/app).
  3. Add a job lint (stage test; after sleep 1, echo lint-ok) with needs: []. When you run it, lint should start before compile finishes.
  4. Put a job audit (stage test) in the long form of needs with job: compile and artifacts: false, and make the script test ! -e bin/app && echo no-artifact. When you run it, audit should succeed (there is no bin/app), and unit should still receive bin/app.
  5. A job integration (stage test) has rules so that it is created only when $RUN_INTEGRATION == "yes", and after sleep 2 it runs echo integration-ok. A job release (stage deploy) has unit and job: integration, optional: true in needs and runs echo release. If you run without the variable, release runs without integration, and if you run with --variable RUN_INTEGRATION=yes, release should start after integration finishes.
  6. Add three more jobs. flaky (test) runs echo flaky-run and then exit 1, but has allow_failure: true. cleanup (deploy) runs, with when: always, echo cleanup, and notify-failure (deploy) runs, with when: on_failure, echo notify. When you run it, flaky should end as a warning, the pipeline should succeed, cleanup should run, and notify-failure should not run. The grader also runs it after turning off flaky's allow_failure in a copy.
  7. Add a job check-config to the stage .pre (without writing it in the stages list). After sleep 2, it fails with test ! -e STOP if there is a STOP file in the repository, and otherwise runs echo config-ok. When you run it, compile should start after check-config finishes. The grader puts a STOP file into a copy and also checks that compile does not run at all when the pre-check fails.

Notes

A stage waits for all of the earlier stage

Make /root/glci-dag a git repository and in .gitlab-ci.yml put stages [build, test, deploy] and four jobs: compile (build; after sleep 3, it creates the file bin/app and uploads bin/ as artifacts), unit (test; test -f bin/app and then echo unit-ok), e2e (test; after sleep 8, echo e2e-ok), and publish-docs (deploy; test -f bin/app and then echo docs-published). Run it with gitlab-ci-local --shell-isolation --no-artifacts-to-source --timestamps and see that publish-docs starts only after the slow e2e has finished.

A job without needs starts only when all jobs of the earlier stage have finished, and it receives the outputs of all the earlier stages. If you add --timestamps, a time is printed in front of each line so you can compare the start (starting shell) and the end (finished in).

Publishing the docs only has to wait for compile

To publish-docs, add needs: [compile]. When you run it, publish-docs should start before e2e finishes, and it should still receive compile's output (bin/app).

needs specifies the jobs to wait for directly. It receives only the outputs of the jobs listed, and the stage order no longer decides the start time. Once you use needs, a job you did not list is neither waited for nor its output received.

needs: [] starts as soon as the pipeline starts

Add a job lint (stage test; after sleep 1, echo lint-ok) with needs: []. When you run it, lint should start before compile finishes.

An empty needs means "wait for nobody". Even if the stage is test, it does not wait for build to finish. If you move forward a check that only needs to look at the source this way, you learn of a failure a few minutes earlier.

Wait for the order but do not receive the outputs

Put a job audit (stage test) in the long form of needs with job: compile and artifacts: false, and make the script test ! -e bin/app && echo no-artifact. When you run it, audit should succeed (there is no bin/app), and unit should still receive bin/app.

The long form of needs lets you turn off, job by job, whether to receive outputs. It keeps a job that needs only the order and not the files from being slowed down by downloading big outputs.

Wait for a job that may or may not exist

A job integration (stage test) has rules so that it is created only when $RUN_INTEGRATION == "yes", and after sleep 2 it runs echo integration-ok. A job release (stage deploy) has unit and job: integration, optional: true in needs and runs echo release. If you run without the variable, release runs without integration, and if you run with --variable RUN_INTEGRATION=yes, release should start after integration finishes.

If you simply write in needs a job that may drop out through rules, GitLab does not create the pipeline itself when that job is absent. optional: true means "wait for it if it exists, and move on if it does not".

A job whose failure is allowed, a job that runs even on failure, and a job that runs only on failure

Add three more jobs. flaky (test) runs echo flaky-run and then exit 1, but has allow_failure: true. cleanup (deploy) runs, with when: always, echo cleanup, and notify-failure (deploy) runs, with when: on_failure, echo notify. When you run it, flaky should end as a warning, the pipeline should succeed, cleanup should run, and notify-failure should not run. The grader also runs it after turning off flaky's allow_failure in a copy.

allow_failure turns a failure into a "warning" so that it does not block later stages. when: on_failure runs only if a job before it has failed, and always runs regardless of the result. An allowed failure does not call on_failure.

A pre-check that runs before all stages

Add a job check-config to the stage .pre (without writing it in the stages list). After sleep 2, it fails with test ! -e STOP if there is a STOP file in the repository, and otherwise runs echo config-ok. When you run it, compile should start after check-config finishes. The grader puts a STOP file into a copy and also checks that compile does not run at all when the pre-check fails.

.pre is a reserved stage that is always at the very front even if you do not write it in stages (and .post is at the very back). You use it when you want to attach a pre-check for the whole pipeline without touching the stage list. If an earlier stage fails, ordinary jobs in later stages are only created and do not run.