TT Lab
Get started
Learn Learning paths Courses

GitLab CI/CD

The build output kept in cache vanished in the deploy job

Continue in TT Lab

Goal

You run the pipeline several times to see in the logs how the cache hits, misses, and is invalidated, confirm artifact exclusion, restriction on the receiving side, and passing values with dotenv, and then build an artifact inspection script.

Why it matters

For artifacts, it is right for a later job to fail if they are missing, and for a cache, it is right for every job to run even if it is missing. If you mix the two, on the day the cache is empty a deployment job quietly ships an empty directory, or artifacts with no expiry fill the repository's capacity. For a cache, everything is about what you make the key from, so if you leave the key as a fixed string, you drag along stale dependencies, and if you let consumers upload as well, contamination spreads. With artifacts, a person cannot check "what was uploaded" every time, so it is safer to automate the inspection.

Steps

  1. Make /root/glci-transfer a git repository (with .gitlab-ci-local/ in .gitignore), and in .gitlab-ci.yml put stages [build, test, deploy] and two jobs. package (build) creates dist/app.tgz and dist/app.js.map and uploads dist/ as artifacts, but excludes dist/*.map and sets expire_in: 1 week. ship (deploy) runs test -f dist/app.tgz && echo has-tgz, test -e dist/app.js.map && echo has-map || echo no-map, and test -d public && echo has-public || echo no-public. When you run it, the ship log must be has-tgz, no-map, and no-public.
  2. Add a job docs (build) that creates public/index.html and uploads public/ (expire_in 2 days) as artifacts. On ship, put dependencies: [package] so that it receives only package's outputs. When you run it, the ship log must still be has-tgz, no-map, and no-public.
  3. In requirements.txt, put the single line requests==2.32.3 and add it to what gets committed. The job deps (build) sets its cache to key: files: [requirements.txt], paths: [.deps/], and policy: pull-push, and after mkdir -p .deps, if .deps/installed exists, prints cache-hit, and if not, cache-miss, and creates that file. If you run it twice in a row, a miss followed by a hit must appear. The grader also changes requirements.txt in a copy and checks that a miss appears again.
  4. Add two more jobs. unit (test) puts a cache with the same key and path as deps at policy: pull, if .deps/installed exists, prints unit-hit, and if not, unit-miss, and then creates .deps/junk. verify-cache (deploy) also receives the same cache with pull and if .deps/junk exists, prints junk-saved, and if not, junk-not-saved. When you run it, unit-hit and junk-not-saved must appear.
  5. Change deps's cache into two lists. The first is key: files: [requirements.txt] with prefix: py added, for .deps/ (pull-push), and the second is, with key: tools-$CI_COMMIT_REF_SLUG, .tools/. Add lines to the deps script that, if .tools/lint exists, print tools-hit, and if not, tools-miss, and create it. Put the same prefix on the keys of unit and verify-cache too. If you run it twice, both caches must be hits in the second run.
  6. Make the job version (build) write the single line VERSION=1.4.2-${CI_COMMIT_SHORT_SHA} to build.env and upload it with artifacts: reports: dotenv: build.env. The job announce (deploy), with needs: [version], runs echo "version=$VERSION". When you run it, the announce log must be version=1.4.2-<짧은 커밋 해시> (the placeholder stands for the short commit hash).
  7. Create /root/glci-transfer/artifact-audit.sh <저장소> (the placeholder stands for the repository). It makes a temporary copy of the repository, commits all files, runs the pipeline, and then, if the uploaded artifacts (under .gitlab-ci-local/artifacts) contain *.map, *.env, *.pem, or a file larger than 1MiB, it ends with the one line BAD <잡>/<경로> … (space-separated, sorted; job and path) and 3, if there are none, with OK and 0, and if the pipeline fails, with ERROR and 1. Reports (under .gitlab-ci-reports/) are left out of the inspection. It leaves nothing in the original repository. The grader checks with this repository (OK) and a copy in which a file that leaks into the artifacts was put.

Notes

Upload only the artifacts you need

Make /root/glci-transfer a git repository (with .gitlab-ci-local/ in .gitignore), and in .gitlab-ci.yml put stages [build, test, deploy] and two jobs. package (build) creates dist/app.tgz and dist/app.js.map and uploads dist/ as artifacts, but excludes dist/*.map and sets expire_in: 1 week. ship (deploy) runs test -f dist/app.tgz && echo has-tgz, test -e dist/app.js.map && echo has-map || echo no-map, and test -d public && echo has-public || echo no-public. When you run it, the ship log must be has-tgz, no-map, and no-public.

artifacts:exclude removes, from what you picked with paths, the files you do not want to upload. It keeps big files that are not needed for deployment, such as source maps and debug symbols, from piling up on every pipeline. expire_in says when to delete, and the actual deletion is done by the server.

A deployment job does not receive even the docs it will not use

Add a job docs (build) that creates public/index.html and uploads public/ (expire_in 2 days) as artifacts. On ship, put dependencies: [package] so that it receives only package's outputs. When you run it, the ship log must still be has-tgz, no-map, and no-public.

A job with neither needs nor dependencies downloads all the outputs of the earlier stages. This is a common reason a deployment job becomes slow from receiving even test reports and docs. dependencies keeps the order by stage and picks only the outputs to receive.

When the lock file changes, the cache changes too

In requirements.txt, put the single line requests==2.32.3 and add it to what gets committed. The job deps (build) sets its cache to key: files: [requirements.txt], paths: [.deps/], and policy: pull-push, and after mkdir -p .deps, if .deps/installed exists, prints cache-hit, and if not, cache-miss, and creates that file. If you run it twice in a row, a miss followed by a hit must appear. The grader also changes requirements.txt in a copy and checks that a miss appears again.

If you give the cache key a list of files, the hash of those files' content becomes the key. When the dependency list changes, the key changes by itself, so you do not drag along a stale cache. The contract of a cache is to write it so that the job can run to the end even without the cache.

Consumers only receive the cache

Add two more jobs. unit (test) puts a cache with the same key and path as deps at policy: pull, if .deps/installed exists, prints unit-hit, and if not, unit-miss, and then creates .deps/junk. verify-cache (deploy) also receives the same cache with pull and if .deps/junk exists, prints junk-saved, and if not, junk-not-saved. When you run it, unit-hit and junk-not-saved must appear.

pull only receives at the start and does not upload at the end. If you set only the one job that creates the cache to pull-push, the time consumers spend recompressing and uploading the same content disappears, and consumers contaminating the cache is also blocked.

Split the key for caches with different lifetimes

Change deps's cache into two lists. The first is key: files: [requirements.txt] with prefix: py added, for .deps/ (pull-push), and the second is, with key: tools-$CI_COMMIT_REF_SLUG, .tools/. Add lines to the deps script that, if .tools/lint exists, print tools-hit, and if not, tools-miss, and create it. Put the same prefix on the keys of unit and verify-cache too. If you run it twice, both caches must be hits in the second run.

You can put several caches on one job. If you mix something that changes along with the lock file, like dependencies, and a tool cache you want to keep separately per branch into one key, a change on one side invalidates the other as well. prefix keeps the name from colliding with other caches that use the same file hash.

Pass a value an earlier job calculated to a later job

Make the job version (build) write the single line VERSION=1.4.2-${CI_COMMIT_SHORT_SHA} to build.env and upload it with artifacts: reports: dotenv: build.env. The job announce (deploy), with needs: [version], runs echo "version=$VERSION". When you run it, the announce log must be version=1.4.2-<짧은 커밋 해시> (the placeholder stands for the short commit hash).

Variables are decided when the pipeline is created, so a value calculated during a run cannot normally be passed to a later job. The dotenv report takes a KEY=VALUE file uploaded as an artifact and puts it into the later job's environment variables. rules has already finished evaluating and cannot see this value.

Catch files that must not go into the artifacts

Create /root/glci-transfer/artifact-audit.sh <저장소> (the placeholder stands for the repository). It makes a temporary copy of the repository, commits all files, runs the pipeline, and then, if the uploaded artifacts (under .gitlab-ci-local/artifacts) contain *.map, *.env, *.pem, or a file larger than 1MiB, it ends with the one line BAD <잡>/<경로> … (space-separated, sorted; job and path) and 3, if there are none, with OK and 0, and if the pipeline fails, with ERROR and 1. Reports (under .gitlab-ci-reports/) are left out of the inspection. It leaves nothing in the original repository. The grader checks with this repository (OK) and a copy in which a file that leaks into the artifacts was put.

gitlab-ci-local puts the artifacts each job uploaded under .gitlab-ci-local/artifacts/<잡이름>/ (the placeholder stands for the job name). A dotenv report is a value passed on purpose, so gitlab-ci-local puts it separately under .gitlab-ci-reports/, and you must leave this place out of the inspection so that a normal configuration does not become BAD. The -size of find can be counted in k units.