Install Order, Permission Boundary and Upgrades
In one sentence
An Operator is an automated administrator that manipulates cluster state on your behalf. That is why if an Operator's service account is compromised, the attacker holds all of that Operator's permissions as they are.
Why it was needed
Most workloads do their work inside their own namespace. When a web application is compromised, the damage is generally limited to that application and its data. An Operator is different. It watches CRs across many namespaces, creates Deployments, reads Secrets, and sometimes even creates RBAC objects. A vulnerability found in a single controller Pod can lead directly to a takeover of the whole cluster.
So the two most important things in running an Operator are not flashy features but installation order and permission boundaries.
How it works
Installation order is a dependency graph. If you do not follow the order, it behaves strangely and quietly.
1. CustomResourceDefinition — 새 타입을 API 에 먼저 등록
2. ServiceAccount / ClusterRole / ClusterRoleBinding — 권한 부여
3. Deployment — 컨트롤러 기동
The reason the CRD must come first is that the controller starts watching that type as soon as it comes up. If the type does not exist, setting up the watch fails and the controller falls into a crash loop. The reason permissions must come first is the same — a controller that starts without permissions gets a 403 from the very first list call. If you deploy with GitOps, you must state this order explicitly with sync waves or dependencies.
The key to permissions is reducing verbs. The most common mistake in a beginner's Operator is verbs: ["*"]. What is actually needed is far narrower.
| What the controller does | Verbs needed | Unnecessarily broad permission |
|---|---|---|
| Read and watch CRs | get, list, watch | create, delete |
| Create/update child resources | get, list, watch, create, update, patch | delete (can be replaced by owner references) |
| Report status | update, patch (status subresource) | update on the whole main resource |
| Record events | create, patch | get, list |
Be especially careful with delete. Most child cleanup can be left to owner references and the garbage collector, so the controller does not need delete permission. Without it, the scenario of a compromised controller mass-deleting resources is blocked at the source.
Another thing that is often forgotten is subresource permissions. Even if you have update permission on webservices, webservices/status is treated as a separate resource and must be granted separately. If you handle finalizers, you also need webservices/finalizers. The symptom "I think I gave the permission, but the controller can't write status" often comes from missing these two.
Leader election prevents duplicate execution. If you run two copies of the controller, the two reconcile the same object at the same time and conflict. Leader election lets only one of several replicas actually do the work, and the leader writes its identity into a Lease object in coordination.k8s.io and renews it periodically. If the leader dies, the lease expires and another replica takes over. The least-privilege sense is needed here too — Lease permission can be limited to that one object with resourceNames.
Do not touch the CRD during an upgrade. Raising the controller image is safe with a rolling update, but deleting an old version of the CRD along with it is dangerous. If that version remains in status.storedVersions, the already stored objects become unreadable. The upgrade procedure is always "with CRDs, only add; remove only after re-storage is done."
An honest note about this lab environment
This lab has no real controller binary, and no way to look inside a Pod (kubectl exec, real logs). So the Operator installation lab deals with the correctness of the manifests. You create and verify the namespace and labels, the service account and the verb set of the ClusterRole, the Deployment's account, leader election arguments, probes, resource limits, and security context, and the leader election Lease. Whether permissions were really applied as intended is checked not from logs but with kubectl auth can-i --as=. This is also the fastest and most reliable way to debug RBAC in real practice.
What you see in the field
First, the one-character binding incident. If the subject name in a ClusterRoleBinding differs by even one character from the actual service account name, the binding is created with no error at all, and only the controller gets a 403. A binding that points to a nonexistent subject is a perfectly valid object.
Second, privilege escalation through a CR. If an Operator trusts a user's CR and creates RBAC objects for them, the user can use the Operator as a detour to gain permissions they could not create directly. The defense has several layers — design the Operator not to create RBAC dynamically, and if it is truly necessary, limit the permissions that can be requested with an allowlist.
Third, permissions are a subject of verification, not documentation. A sentence like "this Operator cannot read Secrets" should not be written on a wiki; it should be confirmed with kubectl auth can-i get secrets --as=... and a no recorded. You must put not only what is allowed but also what must not be allowed in the verification table for it to be evidence that the permission boundary is actually being kept.
What you will do in the next lab
You create a dedicated Operator namespace with labels, document the installation order, and create a least-privilege ClusterRole and binding with no wildcards. You fill the controller Deployment with a dedicated account, leader election, probes, a memory limit, and non-root execution, and create the leader election Lease. Then you check that the CRD versions are preserved while raising the image, diagnose and fix three places in a broken Operator manifest, and finally build a permission verification table that contains both "what is allowed and what must not be allowed," and compare it against the actual responses.