TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

Installing an Operator and Verifying Its Permission Boundary

Continue in TT Lab

Goal

Install a complete set of Operator manifests (namespace, RBAC, controller Deployment, leader election Lease) in the right order, and prove by querying it yourself that its permissions really stay inside the intended boundary.

Why it matters

An Operator can be the workload with the strongest permissions in a cluster. It creates and modifies resources across many namespaces, so if its service account is compromised, the attacker holds those permissions as they are. That is why two things are central. First, installation order is a dependency. The controller starts watching its own type as soon as it comes up, so the CRD must already exist, and it begins with a list call, so the permissions must already exist. If the order is wrong, it shows up as a crash loop or a 403, and it takes time to realize that the cause is not the code but the deployment order. Second, permissions are about reducing verbs. verbs: ["*"] is convenient, but if you leave child cleanup to owner references, you often do not even need delete. And in RBAC, webservices and webservices/status are completely different resources — if you do not know this, you get trapped in the symptom "I gave the permission but it can't write status." Finally, permissions are a subject of verification, not documentation. If you use kubectl auth can-i --as= to confirm both what is allowed and what must not be allowed, that table becomes evidence for the permission boundary.

Steps

Before you start: lab Pods start fresh for every lab, so the cluster state from the previous lab is not there. Steps 2 and 6 require the CRD webservices.apps.labhub.io, and the permission queries in steps 7 and 8 target the crd-lab namespace, so if kubectl get crd webservices.apps.labhub.io returns nothing, reapply the CRD and also run kubectl create ns crd-lab. The CRD must have two versions, v1alpha1 and v1, and the storage version must be v1 (step 6 checks that they are preserved).

  1. Create the namespace labhub-operator and attach the labels app.kubernetes.io/part-of=labhub-operator and pod-security.kubernetes.io/enforce=restricted.
  2. In /root/op/install/out/install-order.txt, write the installation order, one line per stage. Line 1 must mention CustomResourceDefinition, line 2 must mention ServiceAccount and ClusterRole/ClusterRoleBinding, and line 3 must mention Deployment, and do not use the word Deployment or 컨트롤러 (Korean for "controller") in an earlier line (the order check is based on the line where each first appears). The CRD webservices.apps.labhub.io must also exist in the cluster.
  3. In labhub-operator, create a ServiceAccount webservice-controller, and create a ClusterRole webservice-controller with the following rules.
    • apiGroups: ["apps.labhub.io"], resources: ["webservices"], verbs: ["get","list","watch","update","patch"]
    • apiGroups: ["apps.labhub.io"], resources: ["webservices/status","webservices/finalizers"], verbs: ["get","update","patch"]
    • apiGroups: [""], resources: ["events"], verbs: ["create","patch"]
    • apiGroups: ["coordination.k8s.io"], resources: ["leases"], verbs: ["get","list","watch","create","update","patch"] You must not use * anywhere in the verbs or resources. Then create a ClusterRoleBinding webservice-controller and set the subject to kind: ServiceAccount, name: webservice-controller, namespace: labhub-operator.
  4. In labhub-operator, create a Deployment webservice-controller. Set spec.replicas to 1 or 2, spec.template.spec.serviceAccountName to webservice-controller, and the container image to ghcr.io/labhub/webservice-controller:v0.1.0, put --leader-elect=true in args, give the container a livenessProbe and a readinessProbe, and set resources.limits.memory. The Pod's securityContext.runAsNonRoot must be true.
  5. In labhub-operator, create a Lease webservice-controller.labhub.io. spec.holderIdentity is the leader's identity string, spec.leaseDurationSeconds is at least 5, and spec.renewTime is an RFC3339 timestamp with microseconds (date -u +%Y-%m-%dT%H:%M:%S.%6NZ). Then in /root/op/install/out/lease-note.txt, write why it is a problem to have no leader election (the problem of two instances reconciling the same object at the same time).
  6. Raise the Deployment's container image to ghcr.io/labhub/webservice-controller:v0.2.0 and wait until the rollout finishes. The CRD must still have at least 2 versions, and v1 must be in status.storedVersions. Write what you checked before and after the upgrade in /root/op/install/out/upgrade-check.txt, and be sure to include what you checked about storedVersions.
  7. Copy /opt/lab/fixtures/operator/broken-operator.yaml to /root/op/install/fixed-operator.yaml, fix the three wrong places, and apply it. Then in /root/op/install/out/diagnosis.txt, write what was wrong and why, one per line, at least three lines. You must include the binding subject name problem and the missing status subresource permission. After the fix, both kubectl auth can-i update webservices/status -n crd-lab --as=system:serviceaccount:labhub-operator:webservice-controller and kubectl auth can-i list webservices --all-namespaces --as=... must return yes.
  8. Create /root/op/install/out/permissions.json. The top-level key is checks, and each element has verb, resource, and expected (yes/no), and namespace if needed. If you do not write namespace, it is checked across all namespaces. There must be at least 6 entries, at least 2 entries must have expected set to no, and one of them must have resource set to secrets. For example, eight entries are enough: list webservices (yes), watch webservices (yes), update webservices/status in crd-lab (yes), create events in crd-lab (yes), update webservices/finalizers in crd-lab (yes), get secrets in crd-lab (no), delete webservices in crd-lab (no), and create clusterrolebindings (no). The expected of every entry must equal the actual response.

Notes

Create a dedicated namespace for the Operator

Create the namespace labhub-operator and attach the labels app.kubernetes.io/part-of=labhub-operator and pod-security.kubernetes.io/enforce=restricted.

You need a membership label for finding the components all at once, and a policy label that does not allow privileges to Pods. A controller is a workload that needs no privileges at all.

Sort out the installation order

In /root/op/install/out/install-order.txt, write the installation order, one line per stage. Line 1 must mention CustomResourceDefinition, line 2 must mention ServiceAccount and ClusterRole/ClusterRoleBinding, and line 3 must mention Deployment, and do not use the word Deployment or 컨트롤러 (Korean for "controller") in an earlier line (the order check is based on the line where each first appears). The CRD webservices.apps.labhub.io must also exist in the cluster.

The controller starts watching its own type and listing as soon as it comes up. Those two things must be ready first. In the order document, write one line per stage, and do not mention the name of a later stage in an earlier line.

Create least privilege with no wildcards

In labhub-operator, create a ServiceAccount webservice-controller, and create a ClusterRole webservice-controller with the following rules.

In RBAC, a main resource and its subresource are different resources. You must write status and finalizers separately, and events belong to the core group. You must not use an asterisk in either the verbs or the resources.

Write the controller Deployment

In labhub-operator, create a Deployment webservice-controller. Set spec.replicas to 1 or 2, spec.template.spec.serviceAccountName to webservice-controller, and the container image to ghcr.io/labhub/webservice-controller:v0.1.0, put --leader-elect=true in args, give the container a livenessProbe and a readinessProbe, and set resources.limits.memory. The Pod's securityContext.runAsNonRoot must be true.

If it runs as the default service account, the permissions you worked to create are not applied. You need an argument that allows several replicas, two probes that check whether it is alive and whether it is ready to receive traffic, and a memory limit in case the cache grows large.

Create the leader election lease

In labhub-operator, create a Lease webservice-controller.labhub.io. spec.holderIdentity is the leader's identity string, spec.leaseDurationSeconds is at least 5, and spec.renewTime is an RFC3339 timestamp with microseconds (date -u +%Y-%m-%dT%H:%M:%S.%6NZ). Then in /root/op/install/out/lease-note.txt, write why it is a problem to have no leader election (the problem of two instances reconciling the same object at the same time).

A lease holds who the leader is right now, how long it is valid, and when it was last renewed. The renewal time is in a format with microseconds, so it is rejected if the number of digits does not match.

Raise the image and confirm the CRD versions are preserved

Raise the Deployment's container image to ghcr.io/labhub/webservice-controller:v0.2.0 and wait until the rollout finishes. The CRD must still have at least 2 versions, and v1 must be in status.storedVersions. Write what you checked before and after the upgrade in /root/op/install/out/upgrade-check.txt, and be sure to include what you checked about storedVersions.

You can raise the controller image with a rolling update, but you must not casually delete an old version of the CRD. Check the list of versions that have ever been stored first. You also need to confirm that the rollout has finished.

Fix a broken Operator manifest

Copy /opt/lab/fixtures/operator/broken-operator.yaml to /root/op/install/fixed-operator.yaml, fix the three wrong places, and apply it. Then in /root/op/install/out/diagnosis.txt, write what was wrong and why, one per line, at least three lines. You must include the binding subject name problem and the missing status subresource permission. After the fix, both kubectl auth can-i update webservices/status -n crd-lab --as=system:serviceaccount:labhub-operator:webservice-controller and kubectl auth can-i list webservices --all-namespaces --as=... must return yes.

A binding that points to a nonexistent subject is created without an error. And even if you have permission on the main resource, the subresource is separate. Do the checking with permission queries, not logs.

Build a permission boundary verification table

Create /root/op/install/out/permissions.json. The top-level key is checks, and each element has verb, resource, and expected (yes/no), and namespace if needed. If you do not write namespace, it is checked across all namespaces. There must be at least 6 entries, at least 2 entries must have expected set to no, and one of them must have resource set to secrets. For example, eight entries are enough: list webservices (yes), watch webservices (yes), update webservices/status in crd-lab (yes), create events in crd-lab (yes), update webservices/finalizers in crd-lab (yes), get secrets in crd-lab (no), delete webservices in crd-lab (no), and create clusterrolebindings (no). The expected of every entry must equal the actual response.

If you check only what is allowed, you cannot prove the boundary. Include what must not be allowed too, and compare against the actual responses. An entry without a namespace is checked across all namespaces.