TT Lab
Get started
Learn Learning paths Courses

KCSA — Kubernetes Security Associate

Secrets — Better to Revoke Quickly Than to Hide Well

Continue in TT Lab

In one line

The maturity of secret management is measured not by "how well did you hide it" but by "assuming it leaked, how many minutes does it take to revoke and replace it." This one perspective changes all the priorities.

Why this was needed

The standard advice is this — "Don't hardcode it in code; use environment variables, and put .env in .gitignore." This advice blocks exactly one threat. Going into the source repository in plaintext.

The real leak paths are far more numerous. Process memory, crash reports, CI logs, image layers, APM dashboards, and wrongly exposed debug endpoints. Moving to environment variables does not block any of these.

How it works

An environment variable is a delivery method, not a storage location

Loading a .env file and then deleting it is useless. The value is already held by the kernel for each process, and on Linux you can read it like a file.

tr '\0' '\n' < /proc/2841/environ

If you have the same UID or are root, it reads as is. It is no different inside a container — a sidecar in the same Pod, someone who can reach the node, and a debug container attached with kubectl debug --target=app all see the same thing.

It is also good to list the other paths through which environment variables leak.

Where a Kubernetes Secret actually lives

Three criteria for choosing a storage method

"Is it encrypted?" is not a good criterion. Ask these three instead.

  1. How many exposure paths are there?
  2. How many minutes do revocation and replacement take?
  3. Can you tell who read it and when?

Judged by these criteria, the ranking of storage methods is sorted not by "encryption strength" but by "shorter response time during an incident."

The best is to not create static keys

Rotation is a design property, not an operating procedure

This is the most important sentence — if the code assumes only one key, then no matter how good a secret manager you use, the actual replacement does not happen.

If webhook validation looks at only one SECRET, then the moment you change the key, every request signed with the previous key is rejected. There is no way to change the sender and the receiver atomically at the same time, so it becomes effectively impossible to replace, and so nobody rotates.

The solution is separating signing and verification. Sign with the one current key, but attempt verification against the entire set of valid keys. The environment variable name changing from singular to plural (WEBHOOK_SECRET → WEBHOOK_SECRETS) is the marker of this design. With JWT, the kid header + JWKS does the same job — the issuer first publishes the new key in the JWKS, and only after verifiers have fetched it is the signing key changed.

And observability. If you emit a metric tagged with something like key_index recording which key verification succeeded with, you can confirm "have requests signed with the old key really disappeared?" from a graph rather than a guess, and safely delete the old key.

Response time comes before detection

A pre-commit hook is a local setting, so it is bypassed with a single git commit --no-verify, and a new joiner spends days without installing the hook. Real enforcement comes from server-side push protection.

And before investing in detection, there is a question to ask — "assuming this secret has been exposed, how many minutes do revocation and replacement take?" If this number is large, then no matter how densely you do detection, the loss during an incident does not shrink.

What it looks like in the field

The incident response order the author compiled is the definitive version of this perspective. In a situation where a secret has already been committed, the most common reaction is rewriting history, but the order is wrong.

A credential pushed to a public repository is collected by automated scanners within seconds to minutes. And there are things that remain even when you erase the history — forked repositories, the local copies of developers who already cloned, PR references the platform keeps (often still accessible if you know the commit hash), caches of search engines and code search services, and CI caches and artifacts. Moreover, if after cleanup someone pushes from an existing clone as it is, the deleted commit comes back to life.

So the order is revoke → investigate impact → clean up history. For an AWS key, you replace it without downtime in the order: create a new key → deploy → set the old key Inactive → delete, and investigate where that key was used with CloudTrail. Unfamiliar IPs, regions not normally used, and unexpected API calls are the signals for judging a compromise. "History rewriting is not an action that undoes the leak; it is hygiene work that reduces recurrence."

The point that using tools like External Secrets Operator does not exempt you from cluster hygiene is also important. An ExternalSecret manifest has no value, so it is fine to commit to Git, but the operator ultimately materializes the value as a cluster Secret, so the RBAC and encryption-at-rest problems remain as they are.

What you will do in the next lab

In the next lab, you attach PSA labels yourself and confirm that a restricted-violating Pod is rejected. Secrets themselves are covered by quiz rather than by lab, but confirm again that the EncryptionConfiguration you wrote in the earlier module's lab is precisely the answer to the "etcd plaintext storage" problem mentioned here.