TT Lab
Get started
Learn Learning paths Courses

CI/CD Pipelines

With Caching, the Key Is Everything

Continue in TT Lab

One-line summary

There is one rule a cache must keep. Even if you delete the cache, the same artifact must come out. A cache is a device that only reduces time, and the moment it changes the result, it is not a cache but a hidden input. Key design is what makes that rule hold.

Why this is needed

The motive for adding a cache is always speed. But if you build the key carelessly, you lose in both directions. If the key is too tight, nothing matches and the cache might as well not exist, and if the key is too loose, you get builds that pass with stale content. The latter is far worse. Because something that should fail succeeds, and that failure blows up weeks later, after the cache has expired, in an unrelated person's change.

The GitLab documentation states this property very clearly. "Caching is an optimization, but it is not guaranteed to always work. You might need to recreate the cached files for each job that needs them." In other words, it must be a structure that does not need the cache. If a job presupposes the cache's existence, that is not a cache but a dependency.

What to put in the key

A key is a string that writes down "the conditions under which this content may be reused". So everything that affects the result must go in.

The standard form in the GitHub Actions documentation contains exactly this shape.

- uses: actions/cache@v4
  with:
    path: ~/.npm
    key: npm-${{ runner.os }}-${{ hashFiles('**/package-lock.json') }}
    restore-keys: |
      npm-${{ runner.os }}-

GitLab does the same thing with cache:key:files. It is a feature that builds a key tied to the contents of specific files.

What prefix fallback gives and takes away

restore-keys looks for keys that start with a prefix when there is no exactly matching key. The official documentation says it scans in order, and if there are several partial matches, it returns the most recently created cache.

The benefit is clear. Even if you changed one line of the lock file, you do not download all the dependencies anew; you sit on top of the old ones and receive only the difference. The danger is in the same place. Content fetched by prefix fallback has no guarantee of matching the current lock file. So use prefix fallback only for "things that give the same result when recomputed". A download cache is safe, and reusing compiled outputs without checking is dangerous. Skipping the stage only when the key matched exactly, and, on a prefix match, filling in the content and then running the normal procedure again, is the safe form.

One more thing. A cache entry in GitHub Actions cannot have its content changed once it is created. The documentation says that the content of an existing cache cannot be modified and you should create a new cache with a new key. So "I'll just leave the key and fix the content" does not work.

A cache and an artifact are different things

The distinction in the GitLab documentation is the most concise. Use a cache for things like dependencies downloaded from the internet, and use artifacts to pass intermediate results between stages. A cache stays on the runner machine, and an artifact is stored on the server and can be downloaded.

The criterion is one. Is it all right if it disappears? If when it disappears it only takes more time, it is a cache. If when it disappears the next stage cannot run at all, it is an artifact. You occasionally see pipelines that put test reports and coverage results in the cache, but a cache can be evicted, so the evidence from the very moment of failure disappears.

Scope and contamination

A cache is also a trust boundary. If you let any branch save to the cache, then anyone who can push to that branch can tamper with the result of later builds. So platforms keep the scope narrow.

Limits and eviction are also values worth knowing. GitHub Actions defaults to 10 GB per repository, deletes entries that have not been used for more than 7 days, and when space is needed, deletes those with the oldest last access time first. So if you split the key finely to the commit level, the caches push each other out and the hit rate actually drops.

What it looks like in the field

References

What you will do in the next lab

You build a cache manager in shell. You build keys by concatenating the lock file hash, tool version and operating system, and write scripts that save and retrieve with that key. Then you fix one line of the lock file to confirm that the key changes by itself, and make a counterexample of what happens if you let through, as it is, the old content received by prefix fallback. At the end you attach a check that runs again with the cache cleared and compares whether the artifact hash is the same, and a step that counts and reports the hit rate.