CNPE — Cloud Native Platform Engineer
Designing a Platform API and Fencing It
Goal
You design and apply a self-service API called AppClaim as a CRD, make the schema and CEL reject invalid requests, use RBAC to divide what a tenant can and cannot do, and then put a cap on the number issued.
Why it matters
Platform APIs and self-service are a core domain covered in CNPE. And what this domain asks is not whether you can create a CRD but whether you can decide what the API server should be made to reject.
This judgment matters because there are only two places to reject. If the API server rejects, the user sees the reason immediately. If a controller fails later, the user believes it was accepted and waits. Whatever you can catch up front but defer to the back end makes people wander that much more.
This lab has no controller that turns an AppClaim into a real app. But kube-apiserver is real, so CRD registration, OpenAPI validation, CEL evaluation, RBAC decisions, and ResourceQuota admission all actually work. So the rejections and allowances you check here are the same as what happens in a real cluster.
The working directory is /root/cnpe-api, and the tenant namespace is tenant-blue.
Steps
- In
/root/cnpe-api/appclaim-crd.yaml, write the CRD. The group isplatform.labhub.io, the kind isAppClaim, the plural isappclaims, and the scope is Namespaced. There are two versions,v1alpha1andv1, both are served, and the storage version isv1. In bothv1alpha1andv1, include the status subresource and an output column that shows.spec.tier. At the root of both versions, makespecrequired, and within spec maketier,replicas, andmaxReplicasrequired. tier is a string, and the two quantities are integers of 1 or more. After writing it all, apply it. - Add validation to the spec of each of
v1alpha1andv1.tierallows onlybronze,silver, andgold, and a CEL rule preventsreplicasfrom exceedingmaxReplicas. The rejection message must contain the field namemaxReplicas. In both versions, allow the healthy requests bronze (1/1), silver (2/6), and gold (3/3), and check with a server dry-run that these are rejected: a whole spec that is missing, null, or an empty object, a missing required field, a wrong data type, and 0 or negative values. The parentheses show replicas/maxReplicas. - Create the namespace
tenant-blue, and in/root/cnpe-api/claim-checkout.yaml, write the AppClaimcheckout. tier issilver, replicas is 2, and maxReplicas is 6. After you apply it, check that the UID and spec you get when querying through both versions are the same. - Add a transition rule to the spec of each of
v1alpha1andv1to maketierimmutable. The rule must referenceoldSelf, and the rejection message must containtier. A change that leavestieras it is and changes onlyreplicasmust still pass. - In
/root/cnpe-api/tenant-rbac.yaml, write the ServiceAccountblue-dev, the Roleappclaim-author, and the RoleBindingappclaim-author. The role allows get, list, watch, create, update, and patch on AppClaim. The binding must point to that ServiceAccount. - Check with
kubectl auth can-ithat the role has no delete, that it cannot create resourcequotas, roles, and rolebindings, and that it cannot read secrets. The patch and update of AppClaim status and the AppClaim list in kube-system must also be rejected. Compare that a healthy get is yes, and do not interpret a communication or authentication error as no. There must be no*anywhere in the role's verbs, resources, or apiGroups. - Create one more AppClaim,
search, and in/root/cnpe-api/claim-quota.yaml, write the ResourceQuotablue-claims. The cap forcount/appclaims.platform.labhub.iois 2. Check that the quota'sstatus.usedfills in as 2, and if it is empty, diagnose and recheck the CRD Established, API discovery, and the quota controller, and then confirm that a third claim is actually blocked. - In
/root/cnpe-api/api-report.txt, write four lines,storage_version,claims,quota_hard, andtenant_can_delete, in the키=값format (key=value). All four values must be ones you queried directly from the cluster.
Reference
-
python3 /opt/fixtures/cnpe_input_contract.py matrixis a helper that checks 44 requests without storing them. It turns off client-side pre-validation and asks the server directly, so you can tell it apart from a local validation success. If even healthy requests are blocked, it is not a safe API but an unusable API. -
For required values and data types, see the official CRD validation documentation. A rule under spec does not run in its place if the parent spec is missing.
-
Check the capacity, tier, and immutability rules in both served versions. Even if the storage version is v1, it does not substitute for the schema validation of a v1alpha1 request.
-
See the official CRD versioning and the authorization decision.
-
To see whether the server actually rejects, use
kubectl apply --dry-run=server. It goes through the validation path as is while leaving nothing in the cluster. -
The quota key contains dots, which makes it awkward to read with jsonpath. It is better to read it with
kubectl get resourcequota blue-claims -n tenant-blue -o json | jq. -
Right after you apply the CRD, wait for the registration to finish with
kubectl wait --for=condition=Established crd/.... -
One common mistake is creating only the Role and forgetting the RoleBinding. A role is only a definition of permissions, and if you do not attach it, no permissions arise.
-
Another is tying the whole spec together with
self == oldSelf. That is not immutability but freezing, and you would no longer be able to change replicas either.
Decide the CRD skeleton and the two versions
In /root/cnpe-api/appclaim-crd.yaml, write the CRD. The group is platform.labhub.io, the kind is AppClaim, the plural is appclaims, and the scope is Namespaced. There are two versions, v1alpha1 and v1, both are served, and the storage version is v1. In both v1alpha1 and v1, include the status subresource and an output column that shows .spec.tier. At the root of both versions, make spec required, and within spec make tier, replicas, and maxReplicas required. tier is a string, and the two quantities are integers of 1 or more. After writing it all, apply it.
Only one version has storage true. You must leave the old version served as well for it to receive requests for that version. A required inside spec does not make spec itself required, so you also need a required at the root. The status subresource separates status writes into their own path.
Bind the set of values and the relationships between fields
Add validation to the spec of each of v1alpha1 and v1. tier allows only bronze, silver, and gold, and a CEL rule prevents replicas from exceeding maxReplicas. The rejection message must contain the field name maxReplicas. In both versions, allow the healthy requests bronze (1/1), silver (2/6), and gold (3/3), and check with a server dry-run that these are rejected: a whole spec that is missing, null, or an empty object, a missing required field, a wrong data type, and 0 or negative values. The parentheses show replicas/maxReplicas.
The shape of a single value is caught by enum, and the relationship between two fields is caught by CEL. Attach the rule under the spec property with x-kubernetes-validations. The rejection message is the only sentence the user actually reads, so include the field name.
Accept the first self-service request
Create the namespace tenant-blue, and in /root/cnpe-api/claim-checkout.yaml, write the AppClaim checkout. tier is silver, replicas is 2, and maxReplicas is 6. After you apply it, check that the UID and spec you get when querying through both versions are the same.
The two versions do not create different objects. Compare the UID and spec. If you turn on the status subresource, the status in regular create and update requests is ignored. This lab has no AppClaim controller, so do not interpret a successful acceptance as the real app deployment being complete.
Make a field that cannot be changed once set
Add a transition rule to the spec of each of v1alpha1 and v1 to make tier immutable. The rule must reference oldSelf, and the rejection message must contain tier. A change that leaves tier as it is and changes only replicas must still pass.
If a rule references oldSelf, that rule is evaluated only on changes. But if you compare the whole spec with oldSelf, you can no longer change replicas either, and it becomes a freeze. Write the field to bind by name.
Decide what the tenant can do on its own
In /root/cnpe-api/tenant-rbac.yaml, write the ServiceAccount blue-dev, the Role appclaim-author, and the RoleBinding appclaim-author. The role allows get, list, watch, create, update, and patch on AppClaim. The binding must point to that ServiceAccount.
A role is only a definition of permissions, and permissions arise only when you attach a binding. After you create them, be sure to check the decision with kubectl auth can-i. You cannot know the final result by reading the role.
Confirm the closed side with decisions
Check with kubectl auth can-i that the role has no delete, that it cannot create resourcequotas, roles, and rolebindings, and that it cannot read secrets. The patch and update of AppClaim status and the AppClaim list in kube-system must also be rejected. Compare that a healthy get is yes, and do not interpret a communication or authentication error as no. There must be no * anywhere in the role's verbs, resources, or apiGroups.
You saw what was opened in the previous step, so here you look only at the closed side. Delete, quota, role, secret, status change, and reads of other namespaces must all be an explicit no, and there must be no asterisk anywhere in the role. A single asterisk flips all the decisions above.
Put a cap on issuance and see that it blocks
Create one more AppClaim, search, and in /root/cnpe-api/claim-quota.yaml, write the ResourceQuota blue-claims. The cap for count/appclaims.platform.labhub.io is 2. Check that the quota's status.used fills in as 2, and if it is empty, diagnose and recheck the CRD Established, API discovery, and the quota controller, and then confirm that a third claim is actually blocked.
RBAC answers whether you can, and does not answer how many. The quota counts the number. Do not assert a cause merely because status.used is empty. Right after registration, wait for it to be reflected, and if it stays empty, check Established, API discovery, and the controller state. If you delete the quota blindly, you remove the limit.
Record the platform API's current state as values
In /root/cnpe-api/api-report.txt, write four lines, storage_version, claims, quota_hard, and tenant_can_delete, in the 키=값 format (key=value). All four values must be ones you queried directly from the cluster.
Query and write down all four values. The storage version comes from the CRD, the claim count and the cap come from the namespace, and whether delete is possible comes from the permission decision. The grader recomputes the same values and compares them.