TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

What Do Self-Service Fences Actually Block

Continue in TT Lab

One-line summary

A schema covers the content of a request, RBAC covers a principal's permissions, and ResourceQuota covers a namespace's cap. Even an allowed request can be rejected by a quota, and a failed lookup is not evidence of a safe rejection.

Why this was needed

The following is a practice scenario. The Role for blue-dev has only reading and creating AppClaims. Yet it can read Secrets. Before you conclude that the Role file was applied wrongly, you must look at the other bindings attached to the same principal. RBAC permissions add up allow rules. Removing a verb from one Role is not a deny rule that cancels an allow given by another Role.

A RoleBinding can also reference a ClusterRole. In that case the permissions apply within the namespace of that RoleBinding. Do not conclude that it is always a cluster-wide permission just from the name ClusterRole. You must look at which binding connected it with what scope. Official RBAC rules and bindings

How it works

Distinguish permission queries from real requests

kubectl auth can-i queries the authorization decision. An answer of "create is possible" does not mean it will also pass the input schema or the ResourceQuota. Confirm the final behavior with an actual request by that principal or a safe server dry-run. A server dry-run goes through the API's validation path without storing the object, and it is not a way to run all application behavior. Official API dry-run

You can run the following queries in the administrator context of the lab environment. An ordinary user does not automatically gain permission to use --as for an arbitrary principal.

kubectl auth can-i get appclaims.platform.labhub.io \
  -n tenant-blue --as=system:serviceaccount:tenant-blue:blue-dev
kubectl auth can-i patch appclaims.platform.labhub.io \
  --subresource=status -n tenant-blue \
  --as=system:serviceaccount:tenant-blue:blue-dev

If you remove --subresource=status from the second command, it becomes a different question. Putting status in the name position of the resource/name form is also not asking about the subresource permission. Official can-i usage

In the kubectl verified for this lab, an explicit allow appears as yes on stdout with exit code 0, and an explicit denial appears as no with exit code 1. Do not turn an empty stdout, an authentication error, or a connection failure into "it is not yes, so it is no." First confirm a healthy get as a control, and also distinguish the rejection result exactly. If the CLI or the authorization method changes, recheck the output format first.

Separate API usage permission from policy management permission

Request This lab's policy Reason for checking
AppClaim get, list, watch, create, update, patch Allowed Their own team requests, observes, and modifies
AppClaim status patch and update Rejected Separate the principal that writes observed results from the user
Quota create and patch Rejected Users do not modify their own cap
Role and RoleBinding create Rejected Separate the policy management function from the self-service API
Secrets get, AppClaim list in another namespace Rejected Limit credentials and tenant scope
AppClaim delete Rejected The operator reclamation policy chosen in this lab

Here, you must not understand that getting only the permission to create Roles lets you immediately create administrator permissions. Kubernetes has API checks that restrict putting permissions you do not hold into a Role or binding them. escalate and bind are special permissions that can get past the respective checks, so treat them with even more care. This lab does not turn that defense off and also excludes Role management itself from the tenant's duties. Official privilege escalation prevention

A wildcard can widen not only the resources currently needed but also the scope of permissions added later. You cannot judge the safety of the default edit role by its name alone, and it has effects such as Secrets access. Write the user's work in concrete API groups, resources, and verbs, and review the actual bindings too. Official RBAC good practices

A count quota does not measure CPU usage or external cost

count/appclaims.platform.labhub.io: "2" limits the number of objects of this kind. Even if one claim creates 10 external databases, it does not automatically count that cost. A count limit is a tool for preventing control plane objects from growing without bound, and workload volume, cost, and expiration policies must be designed separately. Also, for the quota of an aggregation API that is not CRD-based, you must confirm the responsibility on the extension API server side. Official object count quota

When you diagnose a quota, read three values separately. spec.hard is the desired cap, status.hard is the applied cap, and status.used is the observed usage. The absence of usage does not mean 0. Official ResourceQuota fields

If used is empty in the lab, first check that the CRD is Established, API discovery, the actual AppClaim list, and the quota status. Re-read after giving it time to be reflected, and if it persists, the administrator investigates the controller state. exceeded quota, status unknown, Forbidden, and a connection error are not the same failure. Even if they look like errors with the same name, you must distinguish the cause in the output. Raising the cap or deleting the quota is not a substitute for diagnosis.

Forbidding deletion is not a required Kubernetes rule

This lab decided to block tenant delete and have the operator reclaim. In a real platform, a design is also possible that allows user deletion and has the controller clean up external resources with a finalizer. A finalizer is not executable code but a key that represents a cleanup responsibility. After a delete request, a deletionTimestamp is set, and the object deletion completes only when the controller finishes cleanup and removes the finalizer. If there is no responsible controller, merely adding the key does not cause the cleanup to run. Official finalizer behavior

What it looks like in the field

When we injected a permission query error into an existing LabHub grader, even an empty response passed as "rejection confirmed." After the fix, it requires both a healthy get and an explicit no. For the quota too, it compares the actual count of 2, the cap of 2, and the usage of 2, and checks that the failure of the third request is due to the quota. If you record only the result that it was rejected, you cannot distinguish an authentication outage from a correct policy.

What to do in the next lab

Fill in blue-dev's allow and reject lists and check status and other namespaces separately. When the quota is full, try sending a third claim through both served versions. In the final report, record the values you actually queried, and do not guess and fill in values you could not read.