TT Lab
Get started
Learn Learning paths Courses

KCSA — Kubernetes Security Associate

Why PSP Died and What PSA Does Differently

Continue in TT Lab

In one line

PSP was a structure that "connected policies and users through RBAC," and that was too hard. PSA dropped that connection and simplified it to three namespace labels. It is an exchange: it lost expressiveness and gained operability.

Why this was needed

PodSecurityPolicy worked like this — you create a PSP object, put use permission on that PSP into a ClusterRole, and bind that to the subject that creates Pods (not a person but the SA of the controller that creates Pods). Then at admission time it searches for "among the PSPs this subject can use, is there one that allows this Pod?"

There were several problems.

As a result, many clusters never turned PSP on at all, or finished by giving everyone a permissive PSP. A security feature that is not used is not security. PSP was removed completely in v1.25.

How it works

The three profiles

Profile Character Representative restrictions
privileged Unrestricted None. For system workloads such as CNI and storage drivers
baseline Blocks only known privilege escalations No hostNetwork/hostPID/hostIPC, no privileged, no hostPath, no adding dangerous capabilities
restricted Enforces best practices baseline + runAsNonRoot required, seccompProfile required, all capabilities dropped, allowPrivilegeEscalation=false

To satisfy restricted, the Pod spec needs at least four things — runAsNonRoot: true, allowPrivilegeEscalation: false, capabilities.drop: [ALL], and seccompProfile.type: RuntimeDefault.

The three modes

Mode Behavior When to use
enforce Rejects creation of violating Pods Actual enforcement
audit Records only in the audit log and allows creation Understanding impact
warn Shows a warning to the kubectl user and allows creation Preparing for migration

The label has the form pod-security.kubernetes.io/<모드>: <프로파일> (the placeholders are the mode and the profile), and you can pin the version with <모드>-version (the placeholder is the mode).

labels:
  pod-security.kubernetes.io/enforce: baseline
  pod-security.kubernetes.io/enforce-version: v1.30
  pod-security.kubernetes.io/audit: restricted
  pod-security.kubernetes.io/warn: restricted

Why pinning the version matters. If you do not pin it, it becomes latest, and the moment you upgrade the cluster the policy content may change. This is where the incident comes from in which Pods start being rejected even though you never deployed anything.

The example above is the most common combination in practice — enforce actually blocks with baseline, while audit and warn are set to restricted to collect in advance "what would break if we raised to restricted."

A property you must know: it acts only at admission time

PSA judges only when a Pod is created. Pods that are already running are not evicted even if you raise the label. So right after you raise enforce to restricted, it looks as if nothing happened, and then it blows up at the next rollout or node replacement when Pods fail to come up. There is a time lag between the policy change and the incident.

Exemptions

There are exceptions that cannot be expressed with namespace labels, so exemptions on three axes are placed in the AdmissionConfiguration.

Exemptions are a cluster-wide setting, so you have to restart the apiserver, and you cannot change them with namespace labels. This is also a safeguard — because a namespace administrator cannot exempt themselves.

When PSA is not enough

PSA has only three profiles, so it cannot express things like "allow only images from this registry," "resource limits required on every Pod," or "forbid namespaces without a team label." That is when you add a policy engine.

The judgment the author's blog reached is clear — "Kyverno for simple policies, OPA Gatekeeper for complex cross-resource validation," and the recommended combination is RBAC (basic authorization) + PSA (Pod security baseline) + a policy engine (custom).

Gatekeeper has one hole that an admission webhook misses, so it has a separate Audit Controller. A webhook blocks only new requests, so if you add a policy later, violating resources that were already deployed remain. Audit periodically sweeps everything and finds them. Collecting only violations without blocking, with enforcementAction: dryrun, is the standard procedure for assessing the impact of a new policy.

What it looks like in the field

The author's homelab rebuild log has an incident of exactly the same principle as PSA. Right after installing Cilium, hubble-relay and hubble-ui were Pending, and the reason was 0/1 nodes are available: 1 node(s) had untolerated taint(s). The Deployment did not tolerate the control plane's NoSchedule taint, and in being a type that "is normal behavior but looks like an incident" it has the same nature as a PSA migration. The author's conclusion was the same too — "This is not an error; it is normal behavior."

The standard procedure for a PSA migration is also exactly as the author compiled it — ① audit the current state with a dry run (kubectl label --dry-run=server --overwrite ns --all ...), ② apply warn and audit first, ③ fix the violating workloads, ④ enable enforce, ⑤ apply defaults to new namespaces with AdmissionConfiguration. "A gradual approach is the key, and you never apply enforce all at once."

What to read next

First read one more piece on secret management, and then in the lab you attach PSA labels to a namespace yourself. You confirm that a restricted-violating Pod is rejected and see the rejection message, build a Pod that passes restricted, and finally observe directly that when you raise the level by changing only the label, existing Pods stay alive and only new Pods are rejected.