CKS — Kubernetes Security Specialist
Why Permissions Quietly Grow
In one line
RBAC blocks privilege escalation by default, but the three verbs escalate, bind, and impersonate, and the automatically
mounted ServiceAccount token, get around that line of defense. Hardening means closing these four holes.
Why this was needed
RBAC has a built-in safeguard. For a user to create a Role or RoleBinding, they must already have every permission contained in that Role. Thanks to this rule, the path of "a low-privileged person granting themselves high privileges" is blocked. The problem is that there are verbs that explicitly break through this protection.
| Verb | What it breaks through | Practical response |
|---|---|---|
escalate |
You can add permissions you don't hold to a Role | Platform administrators only, audit log required |
bind |
You can bind a Role with permissions you don't hold | Restrict the permission to create ClusterRoleBindings itself |
impersonate |
You can impersonate another user, group, or SA | Pin the targets with resourceNames and audit |
On top of that, there is the system:masters group. Members of this group bypass the RBAC check entirely.
If a single kubeconfig leaks, that is the end. Keep it only for the break-glass procedure, and normally
nobody should be in this group.
How it works
The second axis is the ServiceAccount token. A Pod runs by default with the default
ServiceAccount of its own namespace, and its token is automatically mounted at /var/run/secrets/kubernetes.io/serviceaccount/token.
It is mounted even if the application doesn't use the Kubernetes API. That is, if a single Pod is
breached, the attacker immediately gets hold of a cluster API credential.
There are two places to turn it off. automountServiceAccountToken: false on the ServiceAccount object,
and the same field in the Pod spec. The Pod-level setting overrides the ServiceAccount level.
So in practice you use both. If you turn off the default SA, the path by which a new workload accidentally receives a token
disappears, and only the workloads that need it get a dedicated SA and turn it on explicitly.
The form of the token has changed too. In the past, a non-expiring token Secret was automatically generated for each SA, and anyone who could read that
Secret obtained a credential valid forever. Now, tokens with a lifetime and an audience are issued
through the TokenRequest API, and they are injected into the Pod through the serviceAccountToken source of a projected volume. A manually created
kubernetes.io/service-account-token Secret can still be created, but since it has no expiry,
you use it only when absolutely necessary.
The third is verification. RBAC is hard to verify by reading it with your eyes. If you add --as=system:serviceaccount:<네임스페이스>:<이름> (the placeholders are the namespace and the name) to kubectl auth can-i,
the API server itself answers what that subject can actually do.
It is the only way to check the outcome rather than the design.
What it looks like in the field
There is a P0 outage that often occurs in organizations that have put a policy engine on top. If OPA Gatekeeper's
ValidatingWebhook is registered with failurePolicy: Fail and all the Gatekeeper Pods go
down, the API server can't get a webhook response and rejects all creation and modification of the target resources.
In trying to enforce policy, cluster deployments come to a complete stop. The emergency fix is to delete the webhook configuration object,
and the root-cause response is to use failurePolicy: Ignore in production but always catch a Gatekeeper outage
with an alert. The fact that a mechanism for tightening permissions can cause an availability incident is a realistic constraint on
hardening design.
The lifetime of credentials was also the issue when the author expanded the control plane from 1 node to 3 in the homelab.
kubeadm init phase upload-certs encrypts the CA key bundle and uploads it as a Secret inside the cluster,
and the certificate-key is the symmetric key that decrypts that ciphertext. This Secret is automatically deleted after 2 hours.
It is a design meant to minimize the time the CA key exists inside the cluster, even in encrypted form, and the same principle
applies as is to SA tokens. For credentials, shortening the lifetime is almost always more effective than hiding them.
What you will do in the next lab
You find the ClusterRoles that have wildcards directly in the cluster and make a list, and you pull out separately the roles that have
dangerous verbs. Then you create a narrowed Role, bind it to a ServiceAccount, and check the result with
auth can-i --as=. Finally, you turn off automatic token mounting on both the Pod and the default ServiceAccount.