Secrets Stay Behind in Logs and Layers
One-line summary
The places where secrets leak in a pipeline are almost fixed. Four places: build logs, image layers and configuration history, repository commit history, and artifacts. And the response to a secret that has leaked is not to delete it but to revoke it and issue a new one.
Why this is needed
Mistakes in handling secrets usually come not from malice but from convenience. Someone printed all the environment variables in order to debug, passing the token as a build argument was the simplest, and someone committed with the value left in a configuration file. Each was something one person did to save a few minutes, but the result is that the range of people who can read that value widens wholesale.
Let us look at the four places one by one.
- Build logs. Logs are usually read widely within an organization, kept for a long time, and searched. One line of a command that prints the value is enough.
- Image layers and configuration history. An image contains not only the file system but also how it was made. Even if you wrote something to a file in the middle of the build and deleted it in a later stage, it remains as it is in the earlier layer.
- Repository commit history. A value that was committed remains in the history even if you delete it in the next commit. It also remains in copies, forks and caches that were taken in the meantime.
- Artifacts. The value gets mixed into configuration files, test reports and dumps that the build produced. Artifacts are often kept longer than logs.
Masking is not a countermeasure but the last line of defense
CI platforms hide known secret values in logs. It helps, but you must not treat it as the countermeasure. The reason is simple. Masking can hide only the same string as one it knows. If the value is altered even slightly, it shows as it is.
- If you print it encoded in base64, it is a different string and is not hidden. The GitHub documentation also states firmly that base64 merely turns binary into text and is no substitute for encryption.
- If it is wedged into JSON or a URL and escaped, it becomes a different string.
- If a multi-line value is printed split by lines, each line differs from the registered value.
- Values derived from a secret (signatures, hashes, token exchange results) are not registered in the first place.
So the order is this. First make it so that it is not printed, and treat masking as the last net that reduces mistakes. GitHub recommends registering even sensitive values that are not secrets themselves with ::add-mask::, but this too carries the same limitation as it is. And in workflows triggered from a forked repository, secrets other than GITHUB_TOKEN are not passed. This is a far stronger protection than masking, and it shows just as it is the principle that dividing the boundary is better than hiding.
Build arguments and runtime secrets
A sentence in Docker's official documentation states the conclusion as it is. "Build arguments and environment variables are inappropriate for passing secrets to a build, because they persist in the final image." Even if you write and then delete the value, it remains in the build record the image contains.
What you use instead is a build-time secret mount.
# 값은 그 RUN 명령이 도는 동안에만 존재하고, 레이어에 남지 않는다
RUN --mount=type=secret,id=npmtoken \
NPM_TOKEN="$(cat /run/secrets/npmtoken)" npm ci
The default mount location is /run/secrets/<id>, and it disappears when that command ends and is not preserved in the image layers.
Runtime secrets are another story. A value needed at run time is given not by the image but by the execution environment. Giving it as an environment variable is convenient, but it is carried along in process lists, child processes, error reports and dumps. Giving it as a file is usually better, and then you decide two things. Permissions (narrow it so that only that process can read it, and look at directory permissions too) and lifetime (when does it disappear, and is it re-read on rotation).
A short lifetime is structurally better
A long-lived credential carries the problem of "we don't know when it leaked". A token that leaked into a log months ago is still valid now. A short-lived credential puts an expiry on the usefulness of a leaked value. Instead of preventing the leak, it cuts the damage window to a few minutes.
For the same reason, grant a narrow permission scope too. If the pipeline needs only to read, give it only read. This is not a measure to prevent a leak but a measure that decides the damage when a leak happens.
The response after a leak is revocation and reissue
This is where people go wrong most often. When a value has gone into a commit, people try to fix the history and erase it. Rewriting history has no effect on copies that have already been cloned, forks, caches, logs or backups. The key point is that you cannot know whether someone read the value in the meantime.
So the order is always the same. First revoke that credential. Issue a new value. Change the places that use it. Only then talk about cleaning the history and preventing recurrence. Erasing it from the history is not the first step, and it is usually the least important step.
What it looks like in the field
- A debug step that prints all the environment variables to investigate a failure is left behind. The preparation for an accident is complete. If needed, print only the key names and not the values.
- If you take apart the image, the token passed as a build argument comes out as it is. The value is already a target for revocation.
- A secret is given as a file with permissions opened wide, so another process in the same Pod can read it.
- A retrospective that concludes "we cleaned the history, so we're fine" when a secret was committed by mistake. If it was not revoked, it is not fine.
References
- Docker build secrets: https://docs.docker.com/build/building/secrets/
- Dockerfile reference: https://docs.docker.com/reference/dockerfile/
- Using secrets in GitHub Actions: https://docs.github.com/en/actions/security-guides/using-secrets-in-github-actions
- GitHub Actions workflow commands: https://docs.github.com/en/actions/using-workflows/workflow-commands-for-github-actions
- GitLab CI/CD variables: https://docs.gitlab.com/ci/variables/
What you will do in the next lab
You reproduce and block the four places where secrets leak one at a time with shell, git and skopeo. First you put in a step that prints environment variables and confirm that the value remains in the log, then you attach a filter that imitates masking and make a counterexample in which you print it converted to base64 and it passes straight through that filter. Next you open the oci-archive in /opt/images with skopeo inspect --raw and see with your own eyes that an image contains not only the file system but also its configuration and history. For the commit history side, you commit a value, delete it, and then pull it out as it is from the old commit. At the end you build a script that checks the permissions and lifetime of secret files and a checklist of the leak response procedure that starts with revocation.