Where Pipelines Cause Incidents
In one sentence
A pipeline is a place that runs arbitrary code holding the strongest privileges in the repository, so most incidents happen not because "the configuration is wrong" but because "someone can steer that execution".
Why this was needed
A CI job runs holding deployment credentials, registry tokens, and cloud keys. And what that job runs is decided by files inside the repository. That is, someone who can put a file into the repository can in effect do anything with those credentials. If you realize this late, you spend years on the phrase "the build server is on the internal network, so it's fine".
Incidents on the supply-chain side have shown this point repeatedly. In the tj-actions/changed-files incident, the attacker did not break into the repository. They only moved the tag of a widely used action to a malicious commit, and the many pipelines that referenced that tag ran the malicious code as it was. A tag is a label a person can move, and the pipelines were trusting that label.
How it works
Defense is placed in four spots.
First, secret values. Put CI variables in the project settings, not in the repository, and mark them as protected variables so that they are passed only to jobs that run on protected branches and protected tags. Masking hides the value in logs, but it is closer to courtesy than defense — if you encode the value in base64 and print it, it passes through masking. The real defense is to keep that value from going into that job in the first place.
Second, merge requests from forks. It is a situation where pipeline configuration written by someone else runs on our runner, so by default you do not give it secret values. Caches also cross a trust boundary. If a fork plants a malicious dependency in the cache, cache poisoning holds, in which later builds use it as it is.
Third, include, which pulls in someone else's configuration. When you bring in a template from another repository, if you set ref to a branch name, the content of our pipeline changes every time that branch changes. You must pin it to a commit SHA for yesterday and today to be the same. The container images a job uses are also pinned to a digest or an immutable tag for the same reason.
Fourth, a revertible deployment. Put the production deployment job on manual approval (when: manual and allow_failure: false), and specify environment to leave a record of what went out where. And if you make the image tag used for deployment the commit SHA, then when an incident happens, the answer to "what was it that went out then" comes within 5 minutes from the registry and the Git log alone. If production has only latest, nobody can answer this question and there is nothing to revert to.
The minimum line to keep in practice
Realistically, it is hard to do everything at once. For a team touching this for the first time, it is better to give an order. First move the deployment credentials to protected variables, next put manual approval on the production deployment, next pin the references of images and include, and lastly touch the pipeline policy for fork merge requests. The first two take a day and prevent most incidents.
One more thing. When the pipeline run time starts to exceed 15 minutes, people stop waiting for the results and go off to do other things. Then failures are found late, and a failure found late is hard to narrow down because several commits have already piled up. Security and speed look like separate topics, but a pipeline that nobody watches is also dead as a gate, so in the end it is the same story.
What you will do in the next lab
You confirm by running the precedence of where variables come in and the environment scope, see how a secret remains in a log without masking, and then build a leak checker and a reference-pinning checker. Then you build a revertible deployment with environment, manual approval, and the commit SHA, and finally a gate that blocks configuration errors before a commit and at the first job of the pipeline. The lab environment has no GitLab server, so protected variables, masking, and CI_JOB_TOKEN are not reproduced, and each lab states that limit. In the quiz after the lab, you tell apart the difference between protected variables and masking, the reason to pin include and image references, and the conditions under which manual approval becomes an approval gate.