CKAD — Kubernetes Application Developer
Closing the Door Is Not Canceling an Order
In one line
Not taking new guests and the kitchen finishing the orders it has already accepted are separate things. The exclusion of traffic in Kubernetes and the shutdown handling of the application must likewise be checked separately.
Why this was needed
After a deployment, the new version opens fine, yet at the moment of deployment a single payment or file-creation request sometimes fails. The indication that a Deployment's replacement has finished does not carry the fate of that request. Whether the new Pod is ready, whether the old Pod has been removed from the targets for new requests, and whether it has finished the requests it was already handling are different questions. If you replace the three questions with one green light, a short status query may pass while a long task is cut off.
This unit does not prove zero downtime for a whole service with several replicas. It puts one Pod behind a Service and follows the life of one request that Pod actually accepted, all the way through. The experiment's scope is kept small not to make the numbers look good, but to identify clearly which connection and which process produced the result. The HTTP response includes the Pod UID injected through the Downward API. A name can be reused, but a UID changes with every creation.
How it works
The readinessProbe expresses whether the Pod is ready to be a target of the Service. The process by which that result is reflected in the EndpointSlice and the traffic path proceeds separately. When a deletion request arrives for a Pod, the ready of a terminating endpoint becomes false. You must not read this as a command to cut existing TCP connections immediately. When a request the server has already accepted and is handling ends is also decided by the server's signal handling and connection behavior.
The conditions of an EndpointSlice include ready, serving, and terminating. terminating signals that it is terminating, and serving can express the actual readiness separately from the terminating state. So an observation where ready is false because it is terminating but serving is true is possible, but you do not always have to look at that combination. Because this server also returns 503 from healthz once drain starts, serving can change through the readiness check that follows. Do not mix the first terminating endpoint and a later endpoint and treat them as one fixed state.
Here we use an ordinary Service that does not turn on publishNotReadyAddresses. Turning this setting on creates an exception in how ready is interpreted. There is also a proxy behavior that, when all available endpoints are terminating, can route to a target where serving and terminating are both true. So ready=false alone does not guarantee that not a single new request arrives. You have to check the app's drain handling and the actual behavior of the traffic path together.
The experiment helper first checks the Service's healthz response. This is because, if you send a work request right after only seeing that the Pod is Ready, a connection error from a Service path that is not yet ready can be mistaken for a termination effect. Next, it sends a work request, confirms that the server's active is 1, and then deletes. The observation order is connection ready, server acceptance, deletion request, response or disconnection. If there is no evidence of acceptance in this, you cannot reach the conclusion that "the request being handled was cut off" either.
What it looks like in the field
With an API gateway, a service mesh, and an external load balancer, there are more layers that pick traffic. You cannot carry a result obtained from a single k3s Service over as a guarantee for other paths. In real operation, you must observe separately which layer stops new connections and which layer keeps existing ones. Look not only at the success rate but also at request latency, the response body, and the side effects that retries created.
How you count failures also matters. A curl transfer success and an HTTP 200 are not the same condition. You can receive HTTP 500 and have the connection be fine, so it can be counted as a success. Conversely, you cannot conclude that a connection-ready timeout was the app's forced termination either. This lab records the normal completion body, the connection-drop exception, and the Pod exit code separately to distinguish different failures.
What you will do in the next lab
Compare a server that exits immediately with a server that waits for existing requests, and record the conditions of the first terminating endpoint. Explain, with the response and the UID, that the single value ready=false does not stand in for the result of a request that was already accepted.
Official documentation: Observing Pod and endpoint termination, EndpointSlice.