Building the Pipeline Execution Model From Configuration
Goal
You stack up a single .gitlab-ci.yml over eight steps and build by hand the execution model of GitLab CI — stage order, the DAG that needs makes, whether a job is created as decided by rules, and the difference between artifacts and cache. At the end, you write the interpreter that calculates that model yourself.
Why it matters
This Pod has neither a GitLab server nor a GitLab Runner. So this lab does not pretend that a pipeline is actually running. Instead, it covers only what you can honestly learn even without a runner — the configuration language and its execution model. In practice, the moments when you get stuck because of a pipeline are mostly not because you don't know the commands but because you cannot explain why this job was not created and why that job is still waiting. The answers all come from the structure of this file. Grading also does not skim the file with eyes but reads it with a YAML parser and checks the structure — because a configuration in which one wrong space of indentation leaves a job hanging under the wrong key looks fine to human eyes and does not actually run.
Steps
- Create
/root/glci/.gitlab-ci.yml. Under the top-levelstages, putbuild,test, anddeployin this order, and put the jobbuild-appwithstage: buildand write at least onescript. - Add more jobs to fill all three stages.
lintandunit-testarestage: test, anddeploy-stagingisstage: deploy. Every job must have ascript, and everystagevalue must be in thestageslist. - Make a hidden job called
.python-basewithimageandbefore_script, and have the three jobsbuild-app,lint, andunit-testinherit it withextends: .python-base. The inheriting jobs do not writeimageagain. In the top-leveldefault, put a defaultimage. - Add
needsto make a DAG.lintgetsneeds: [],unit-testgets, inneeds,build-app, anddeploy-staginggets, inneeds,unit-test.build-appgets noneeds. - To
deploy-staging, addrulesso that when$CI_COMMIT_BRANCH == "main", it haswhen: on_success, and close the last entry with awhen: neverwith no condition. Create a new jobdeploy-prodwithstage: deploy, in itsneeds, writedeploy-staging, set it towhen: manualandallow_failure: falseunder the same branch condition, and likewise close it with awhen: neverwith no condition. Do not useonlyorexcept. - To
build-app, addartifactsand writepathsandexpire_in. Forunit-test, changeneedsto the long form withjob: build-appandartifacts: true, and fordeploy-staging, setneedstoartifacts: false. - Make
/root/glci/requirements.txta file with content. Tobuild-appandunit-test, addcache, and makekeypoint, withfiles, torequirements.txt.build-appispolicy: pull-pushandunit-testispolicy: pull. Keepcache.pathsfrom overlappingartifacts.paths. - Create
/root/glci/plan.py. It reads the configuration file given as an argument and prints{"stages": [...], "waves": [[...], ...]}as JSON to standard output. If there isneeds, it waits for only that list, and if not, it waits for all jobs in the earlier stages (the defaultstageistest). Keys starting with a dot and reserved keys (stages,variables,default,include,workflow) are not jobs. Sort each wave by name. If no job can start, it prints the wordcycleand ends with a non-zero exit code. Finally, save the output of running it on your own configuration to/root/glci/plan.json.
Notes
- If you build the habit of checking with a parser, incidents are halved: print the top-level key list with
python3 -c "import yaml,sys;print(list(yaml.safe_load(open(sys.argv[1]))))" /root/glci/.gitlab-ci.yml, and you can see at once whether a job has disappeared. needs: []and having noneedskey at all have opposite meanings. The former waits for nothing, and the latter waits for everything in the earlier stages.rulesapplies only the first matching entry and stops. Always put awhen: neverwith no condition at the very end of the list.- Common mistakes: leaving out the dot in a fragment name and making it a job that runs, inheriting with
extendsand then writingimageagain, making the cache path and the artifacts path the same, and sending the interpreter's diagnostic messages to standard output and breaking the JSON.
Stage order and the first job
Create /root/glci/.gitlab-ci.yml. Under the top-level stages, put build, test, and deploy in this order, and put the job build-app with stage: build and write at least one script.
Create /root/glci/.gitlab-ci.yml and put a stages list at the top level. The order of this list is the default execution order. Under it, put a top-level key called build-app, and indent stage and script inside it. Grading reads with a YAML parser, not grep, so even one wrong space of indentation makes the job a child of another key and it fails.
Fill all three stages
Add more jobs to fill all three stages. lint and unit-test are stage: test, and deploy-staging is stage: deploy. Every job must have a script, and every stage value must be in the stages list.
Add lint and unit-test to the test stage and deploy-staging to the deploy stage. If there are two jobs in the same stage, those two run in parallel with each other. If you use a name that is not in the stages list as stage, GitLab rejects that configuration entirely, so check the spelling. Every job must have a script.
Strip out duplication with hidden jobs and extends
Make a hidden job called .python-base with image and before_script, and have the three jobs build-app, lint, and unit-test inherit it with extends: .python-base. The inheriting jobs do not write image again. In the top-level default, put a default image.
A key whose name starts with a dot is a fragment that does not run. In .python-base, put image and before_script, and have the three jobs build-app, lint, and unit-test inherit it with extends. Once inherited, do not write image again in the job. For jobs that do not use the fragment, put a default image in the top-level default.
Cross the stage walls with needs
Add needs to make a DAG. lint gets needs: [], unit-test gets, in needs, build-app, and deploy-staging gets, in needs, unit-test. build-app gets no needs.
Make unit-test, with needs, wait only for build-app, and make deploy-staging wait only for unit-test. Give lint needs: [] — a missing key and an empty list have opposite meanings, so only with an empty list does it wait for nothing in the earlier stages and start at the very front. build-app gets no needs.
Apply branch conditions and manual approval with rules
To deploy-staging, add rules so that when $CI_COMMIT_BRANCH == "main", it has when: on_success, and close the last entry with a when: never with no condition. Create a new job deploy-prod with stage: deploy, in its needs, write deploy-staging, set it to when: manual and allow_failure: false under the same branch condition, and likewise close it with a when: never with no condition. Do not use only or except.
For deploy-staging, when $CI_COMMIT_BRANCH == "main", use when: on_success, and for deploy-prod, under the same condition, use when: manual. The last entry of both jobs must be a when: never with no condition — rules applies only the first match, so if you put an entry with no condition in the middle, everything below it dies. Do not use only/except. deploy-prod is in the deploy stage and waits for deploy-staging.
Decide artifacts and who receives them
To build-app, add artifacts and write paths and expire_in. For unit-test, change needs to the long form with job: build-app and artifacts: true, and for deploy-staging, set needs to artifacts: false.
To build-app, add artifacts, write the directories to hand over in paths, and write expire_in as well. Then change needs to the long form (- job: 이름 / artifacts: true|false, where the placeholder stands for the job name), so that only unit-test, which actually uses the output, receives with true, and deploy-staging, which only waits for order, turns it off with false. If you download even for jobs that don't use it, the pipeline quietly becomes slow.
Make the cache key the hash of the lock file
Make /root/glci/requirements.txt a file with content. To build-app and unit-test, add cache, and make key point, with files, to requirements.txt. build-app is policy: pull-push and unit-test is policy: pull. Keep cache.paths from overlapping artifacts.paths.
First make /root/glci/requirements.txt — there must be a real file for the cache key to hash. Then, to build-app and unit-test, add cache, but make key point to that file in the files form rather than a fixed string. build-app, which creates the cache, is policy: pull-push, and unit-test, which only reads, is policy: pull. cache.paths must not overlap artifacts.paths.
Calculate the execution order with a pipeline interpreter
Create /root/glci/plan.py. It reads the configuration file given as an argument and prints {"stages": [...], "waves": [[...], ...]} as JSON to standard output. If there is needs, it waits for only that list, and if not, it waits for all jobs in the earlier stages (the default stage is test). Keys starting with a dot and reserved keys (stages, variables, default, include, workflow) are not jobs. Sort each wave by name. If no job can start, it prints the word cycle and ends with a non-zero exit code. Finally, save the output of running it on your own configuration to /root/glci/plan.json.
When called as /root/glci/plan.py <설정파일> (the placeholder stands for the configuration file), it must print {"stages": [...], "waves": [[...], ...]} to standard output. There are three rules for calculating waves — if there is needs, it waits for only that list, and if not, it waits for all jobs in the earlier stages (the default stage is test), and keys starting with a dot and reserved keys are not jobs. Sort each wave by name. If it gets into a state where nobody can start, it is a cycle, so print cycle and end with a non-zero code. Send diagnostics to standard error so that standard output remains JSON. At the end, save the result of running on your own configuration as /root/glci/plan.json.