artifacts and cache Are Different Things
In one sentence
artifacts are outputs that cross from job to job, so it is right for a later job to fail if they are missing, while cache is a temporary asset that saves time between runs, so it is right for every job to just run even if it is missing.
Why this was needed
Both look like "if you specify a directory, it is stored somewhere and comes back later". So at first you use whichever, and before long two kinds of incidents happen.
A team that used a cache as an output finds one day that a deployment job deploys an empty directory. A cache is an optimization, not a guarantee, so if the runner changes or it expires, it is simply empty, and then the pipeline does not fail but quietly succeeds with nothing in it. Conversely, a team that used outputs like a cache gets a repository capacity warning. Artifacts with no expiry set pile up with every commit.
How it works
artifacts upload the specified paths to the GitLab server when a job ends. And later jobs download them when they start. What matters here is who downloads, and since the default is "everything from earlier stages", if you configure nothing, a deployment job downloads even the test reports it will never use. This is a common reason pipelines become slow.
The way to narrow the receiving side is the long form of needs. If you write needs not as a list of strings but in the form {job: build-app, artifacts: true}, you can turn off, job by job, whether to receive. For a job that only waits for order and does not need the files, state artifacts: false explicitly. And expire_in is effectively not optional but required. Keep only the release outputs that are candidates for rollback for a long time, and set the rest short, to a few days.
cache is different. When a job starts, it downloads and unpacks the archive matching the key, and when it ends, it uploads it again. It is all optimization, so the job goes on even if it fails. So the core of cache design is not what you put in but what you make the key from.
If you leave the key as a fixed string, you keep using the same key even when the dependencies change and drag along a stale cache. So you put the hash of the lock file in the key. If you write it like key: {files: [requirements.txt]}, the key changes by itself when that file changes and the cache is invalidated automatically. If you add policy to this, you save once more. If you set only the one job that builds and uploads the cache to pull-push and set the other consumers to pull, the time the consumers spend recompressing and uploading the same content every time they finish disappears entirely.
Finally, the two lists must not overlap. If you manage the same directory as both artifacts and cache, nobody knows which one is newer, and that confusion is hard to reproduce.
What you see in the field
The most dangerous of the cache-related incidents is about trust, not performance. If a merge request from a fork plants a malicious dependency in the cache, later builds take that cache and use it as it is. Because of this problem, called cache poisoning, caches that cross a trust boundary get separate scopes.
Capacity is also a realistic constraint. Cache storage has a limit, and when it is exceeded, the oldest are pushed out first. So if you slice the key too finely so as to create a new cache for every commit, caches push each other out and the hit rate actually drops. The key should be set by the unit of reuse.
What you will do in the next lab
You stack up everything you have read so far into a single configuration file yourself. Starting from stages and a first job, you add, in eight steps, hidden jobs and extends, a DAG made with needs, the branch conditions and manual approval of rules, and then artifacts and cache. Grading does not skim the file with eyes but reads it with a YAML parser and checks the structure. In the last step, you write a pipeline interpreter yourself that calculates, even for a configuration you are seeing for the first time, in which wave each job starts and catches needs cycles.