TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

Self-service APIs that own changes and retirement

Continue in TT Lab

One-line summary

Good self-service keeps converging the difference between the request and the actual state even after creation, distinguishes a deletion request from the actual completion of reclamation, and preserves other teams' resources.

Why this was needed

In a platform demo, an app gets applause once it is created. In production, what comes after is longer. You increase the replica count, apply new settings, and delete environments you no longer use. The user sees the response "changed," but some Pods may still respond with the previous settings. The deletion request succeeded, yet a resource that incurs cost may remain. If you automate only creation and leave the rest to people, the operational burden comes back under a different name.

Consider a situation where a developer edited the child Deployment directly. The App request says 2, and if you change the Deployment to 3, which is the source of truth? In this lab, the App is the source. The controller observes the difference and returns the child resource to 2. The reversion itself is not a bug but the result of the declared contract. A change you want to keep must be made on the source request.

How it works

Change verification starts with whether it is a request with the same UID as before. Kubernetes names can be reused. If you delete parcel and create it anew with the same name, the URL and name look the same but the UID is different. This is why the success record of the previous app must not be used as evidence for the new app. If you also link the child Deployment's UID, you can tell apart a workaround that deleted and recreated the resource.

Next is the generation. The Deployment's metadata.generation expresses a change in the desired configuration, and status.observedGeneration shows the generation the controller has observed. A state that was Ready in the previous generation can still be visible briefly right after a new declaration. So you check together whether the observed generation is the current generation, whether the number of new replicas and available replicas equals the target, whether old Pods that are terminating remain, and whether the current Pod actually returns the preview response.

If you change the child resource to 3 as an experiment and submit only the final result of 2, you cannot tell whether you actually created the difference. It would be the same result even if it had been 2 from the start. The experiment helper in this lab keeps the patch response of the 3 state that the API stored. By checking that it returned to 2 in a generation later than that one, it connects the before and after of the change. A single state and a flow of events are different kinds of evidence.

Deletion also has stages. An API server that receives a delete request can mark a deletionTimestamp. If a finalizer remains, the object does not disappear right away. This marker is a device to give the responsible controller a chance to finish its cleanup procedure. If you remove all finalizers during incident response, the object in front of you may disappear, but external resources the owner should be responsible for can be left behind.

In the lab, you insert one educational hold marker and observe the state of being deleted. This marker only delays the absence of the parent request and does not guarantee that all child resources are alive. The controller may reclaim the child resources first, so you must look at the actual state each time. When you release it, remove only that one marker, and compare the UID and resourceVersion so that you do not overwrite changes made in between.

What it looks like in the field

"There is no query result" also has two meanings. One is when the list API responded normally and there is no target, and the other is when you could not obtain the list because of an authentication or communication problem. If you turn the latter into an empty list, the deletion checker reports an outage as a success. The lab acknowledges absence only when the list queries for App, Deployment, ReplicaSet, Service, and Pod each succeed and there is no target.

You must also check the scope. If you clean up team-a and delete team-b as well, my app is certainly gone. But that is not a successful cleanup. You compare whether the UID and the real response of the other team's sentinel are unchanged, to check that only the necessary resources were reclaimed. This lab deals with separating team permissions inside a personal VM, and does not claim that namespaces alone complete all network and resource isolation.

What to do in the next lab

You change the same request to preview and 2 replicas, and observe a manual change to a child resource reverting to the source. Next you record the being-deleted state and the actual reclamation separately and check that the other team was preserved. At the end, you write a small diagnostic tool that leaves observation failures and contradictions as unknown. The goal is the habit of not presuming success from missing information.

Official documentation