TT Lab
Get started
Learn Learning paths Courses

CKAD — Kubernetes Application Developer

Grace Time Is a Budget: preStop and TERM

Continue in TT Lab

In one line

preStop is not bonus time outside the termination grace period. You have to design together the hook of the Pod being terminated, the application's signal handling, and the remaining working time.

Why this was needed

It is easy to think that if you attached readiness and waited a moment in preStop, you are safe no matter what. But if the process exits right after the hook, requests already received can be cut off, and even if the process waits for requests, it cannot finish if the termination grace period is used up first. The purpose of this lab is to fill the gap between "we put the setting in" and "we got the result we wanted."

Which Pod the setting is in also matters. A preStop newly added to a Deployment template is a setting for the Pods created from that template. If you interpret it as having been applied retroactively to the old Pods that are being replaced and taken down now, the experiment's comparison is wrong. You have to read and preserve the UID of the target Pod and the lifecycle in spec.containers just before deletion. A report that captured only the new template cannot tell you what was actually run.

How it works

In the ordinary TERM termination flow, the kubelet runs the preStop hook and, after the hook ends, sends the termination signal to the container process. The termination grace calculation starts before the hook runs, so the hook and the cleanup after it must fit in the same budget. If the hook used 3 seconds, those 3 seconds are not credited separately to the application. Putting in an arbitrarily long sleep can give time to excluding new requests, but it also consumes the termination budget.

People also get the time sum wrong here. An HTTP task that has already started can keep progressing while the hook runs. So even if the total task time is 10 seconds and the hook is 3 seconds, there are not always 13 seconds of remaining work. When designing, work out the hook run time and the upper bound of the processing and cleanup time that actually remains after TERM, and when observing, record the times of acceptance, deletion, hook, TERM, and completion separately. If there is no upper bound on the task, only raising the grace cannot guarantee a safe shutdown.

In this server's graceful mode, on receiving TERM it rejects new work requests, waits until active reaches 0, and then exits. The immediate mode, for comparison, ends at once with exit code 17. The code 17 itself is not a standard termination meaning in Kubernetes. It is an observation marker the lab server chose on purpose. In normal completion, the response body is sent to the end and exit code 0 is observed.

Even with graceful, a container can be forcibly terminated if the grace is insufficient. In the real-VM probe run earlier this time, 137 was observed when the grace was insufficient. However, in the field, do not conclude from 137 alone that it was the termination grace. Other causes such as OOM also have to be investigated. The key is to connect together the termination reason from the Pod watch, the actual grace setting, the times before and after the signal, and the request result.

What it looks like in the field

The wall-clock time when kubectl delete ran and the time the kubelet began handling the termination are not exactly the same. There are also delays in API processing, observation delivery, and runtime work. That is why grading like "the grace is 5 seconds, so the connection must be cut exactly 5 seconds after the delete command" can be wrong even in a healthy environment. This lab does not make you hit fractions of a second; it checks the result of the same accepted request and the termination state.

Querying the Pod again at grading time is also not enough. A Pod that has already been deleted no longer exists, and a terminated container may have been reclaimed by the runtime. In the probe's first implementation, trying to check after termination lost the evidence. That is why the Pod watch is started before the request is sent, and the final container state and the DELETED event are preserved. Regrading that reads the observation file does not recreate the Pod. This is because quietly running the same experiment once more would not be inspecting the previous result.

What you will do in the next lab

With the normal termination configuration, shorten the grace to cut the request, then fix the grace again and check whether a task of the same length completes. Compare the result you expected after reading the reading with the actual observation, but do not expand a single response from one particular run into a zero-downtime guarantee for all production traffic. When you attach an automatic retry after a request failure, you also have to consider the business side effects, like the Job idempotency in the earlier unit.

Official documentation: Container lifecycle hooks, Pod termination flow.