One slow e2e job made docs publishing wait eight seconds
Goal
You actually run a pipeline that has slow jobs mixed in, measure how stages and needs change when each job starts, and confirm from the run results how artifact passing, waiting for conditional jobs, and allowing failures behave.
Why it matters
The stage approach is easy to understand, but the slowest job holds up the start of all later jobs. If you make jobs wait only for what they need with needs, the pipeline gets faster, but in exchange you must write exactly, job by job, what they wait for and what they receive. If you wait for a job that may drop out through rules, the pipeline is not created at all, and if you allow failures too broadly, a real incident passes by in green. These choices become clear when you look at times and results rather than when you read the documentation.
Steps
- Make
/root/glci-daga git repository and in.gitlab-ci.ymlput stages[build, test, deploy]and four jobs:compile(build; aftersleep 3, it creates the filebin/appand uploadsbin/as artifacts),unit(test;test -f bin/appand thenecho unit-ok),e2e(test; aftersleep 8,echo e2e-ok), andpublish-docs(deploy;test -f bin/appand thenecho docs-published). Run it withgitlab-ci-local --shell-isolation --no-artifacts-to-source --timestampsand see that publish-docs starts only after the slow e2e has finished. - To
publish-docs, addneeds: [compile]. When you run it, publish-docs should start before e2e finishes, and it should still receive compile's output (bin/app). - Add a job
lint(stage test; aftersleep 1,echo lint-ok) withneeds: []. When you run it, lint should start before compile finishes. - Put a job
audit(stage test) in the long form ofneedswithjob: compileandartifacts: false, and make the scripttest ! -e bin/app && echo no-artifact. When you run it, audit should succeed (there is no bin/app), and unit should still receive bin/app. - A job
integration(stage test) has rules so that it is created only when$RUN_INTEGRATION == "yes", and aftersleep 2it runsecho integration-ok. A jobrelease(stage deploy) hasunitandjob: integration, optional: truein needs and runsecho release. If you run without the variable, release runs without integration, and if you run with--variable RUN_INTEGRATION=yes, release should start after integration finishes. - Add three more jobs.
flaky(test) runsecho flaky-runand thenexit 1, but hasallow_failure: true.cleanup(deploy) runs, withwhen: always,echo cleanup, andnotify-failure(deploy) runs, withwhen: on_failure,echo notify. When you run it, flaky should end as a warning, the pipeline should succeed, cleanup should run, and notify-failure should not run. The grader also runs it after turning off flaky's allow_failure in a copy. - Add a job
check-configto the stage.pre(without writing it in the stages list). Aftersleep 2, it fails withtest ! -e STOPif there is a STOP file in the repository, and otherwise runsecho config-ok. When you run it, compile should start after check-config finishes. The grader puts a STOP file into a copy and also checks that compile does not run at all when the pre-check fails.
Notes
- This VM has no GitLab server or runner, and gitlab-ci-local 4.75.1 interprets .gitlab-ci.yml by the same rules as GitLab and runs jobs with the shell. If you write
image:, it tries to run with Docker, so do not use it. Protected variables, masking, CI_JOB_TOKEN, runner tags, and merge request pipeline creation are server features and are not reproduced here. - Run: at the repository root,
gitlab-ci-local --shell-isolation --no-artifacts-to-source(a separate working directory per job, artifacts not written back to the repository), job list:gitlab-ci-local --list-csv-all, interpreted configuration:gitlab-ci-local --preview. gitlab-ci-local passes only files tracked by git to jobs, sogit adda file after creating it. The grader copies the repository, commits all files, and runs it again with the same tool. - If the job that needs points to is in the same stage or in a stage that has not finished yet, gitlab-ci-local waits until that stage ends (measured). GitLab waits only for the job pointed to, so on a real server it starts earlier. The time comparison in this lab uses only cases where the two tools give the same result.
- gitlab-ci-local does not reject a job that dropped out through rules even if you write it in needs without optional. GitLab does not create the pipeline in that case and gives an error.
- gitlab-ci-local 4.75.1 handles retry, timeout, allow_failure:exit_codes, and the .post stage differently from GitLab (measured: it does not retry, it keeps running even past the time limit, it proceeds to later stages even when a job fails with an exit code that was not allowed, and .post jobs are only in the list and are not run). So this lab does not cover them.
- needs · CI/CD YAML syntax reference (allow_failure, when) · Job artifacts · Pipeline efficiency
A stage waits for all of the earlier stage
Make /root/glci-dag a git repository and in .gitlab-ci.yml put stages [build, test, deploy] and four jobs: compile (build; after sleep 3, it creates the file bin/app and uploads bin/ as artifacts), unit (test; test -f bin/app and then echo unit-ok), e2e (test; after sleep 8, echo e2e-ok), and publish-docs (deploy; test -f bin/app and then echo docs-published). Run it with gitlab-ci-local --shell-isolation --no-artifacts-to-source --timestamps and see that publish-docs starts only after the slow e2e has finished.
A job without needs starts only when all jobs of the earlier stage have finished, and it receives the outputs of all the earlier stages. If you add --timestamps, a time is printed in front of each line so you can compare the start (starting shell) and the end (finished in).
Publishing the docs only has to wait for compile
To publish-docs, add needs: [compile]. When you run it, publish-docs should start before e2e finishes, and it should still receive compile's output (bin/app).
needs specifies the jobs to wait for directly. It receives only the outputs of the jobs listed, and the stage order no longer decides the start time. Once you use needs, a job you did not list is neither waited for nor its output received.
needs: [] starts as soon as the pipeline starts
Add a job lint (stage test; after sleep 1, echo lint-ok) with needs: []. When you run it, lint should start before compile finishes.
An empty needs means "wait for nobody". Even if the stage is test, it does not wait for build to finish. If you move forward a check that only needs to look at the source this way, you learn of a failure a few minutes earlier.
Wait for the order but do not receive the outputs
Put a job audit (stage test) in the long form of needs with job: compile and artifacts: false, and make the script test ! -e bin/app && echo no-artifact. When you run it, audit should succeed (there is no bin/app), and unit should still receive bin/app.
The long form of needs lets you turn off, job by job, whether to receive outputs. It keeps a job that needs only the order and not the files from being slowed down by downloading big outputs.
Wait for a job that may or may not exist
A job integration (stage test) has rules so that it is created only when $RUN_INTEGRATION == "yes", and after sleep 2 it runs echo integration-ok. A job release (stage deploy) has unit and job: integration, optional: true in needs and runs echo release. If you run without the variable, release runs without integration, and if you run with --variable RUN_INTEGRATION=yes, release should start after integration finishes.
If you simply write in needs a job that may drop out through rules, GitLab does not create the pipeline itself when that job is absent. optional: true means "wait for it if it exists, and move on if it does not".
A job whose failure is allowed, a job that runs even on failure, and a job that runs only on failure
Add three more jobs. flaky (test) runs echo flaky-run and then exit 1, but has allow_failure: true. cleanup (deploy) runs, with when: always, echo cleanup, and notify-failure (deploy) runs, with when: on_failure, echo notify. When you run it, flaky should end as a warning, the pipeline should succeed, cleanup should run, and notify-failure should not run. The grader also runs it after turning off flaky's allow_failure in a copy.
allow_failure turns a failure into a "warning" so that it does not block later stages. when: on_failure runs only if a job before it has failed, and always runs regardless of the result. An allowed failure does not call on_failure.
A pre-check that runs before all stages
Add a job check-config to the stage .pre (without writing it in the stages list). After sleep 2, it fails with test ! -e STOP if there is a STOP file in the repository, and otherwise runs echo config-ok. When you run it, compile should start after check-config finishes. The grader puts a STOP file into a copy and also checks that compile does not run at all when the pre-check fails.
.pre is a reserved stage that is always at the very front even if you do not write it in stages (and .post is at the very back). You use it when you want to attach a pre-check for the whole pipeline without touching the stage list. If an earlier stage fails, ordinary jobs in later stages are only created and do not run.