TT Lab
Get started
Learn Learning paths Courses

CKAD — Kubernetes Application Developer

Between Configuration and Behaviour

Continue in TT Lab

In one line

Probes, Jobs, QoS, and configuration injection are all things the kubelet does. Without a kubelet you can practice writing YAML, but you can never see what it actually does.

Why you need to see it in action

CKAD asks "how does an application live on a cluster?" Probes, Jobs, QoS, and configuration injection are all part of that story.

But the fake cluster that the earlier modules run on has no kubelet, so none of these actually happen. Attach a probe and it does not run, create a Job and it does not complete, and a ConfigMap is not mounted.

How the three probes relate

Probe What it asks If it fails
startupProbe Has startup finished? Keeps postponing the other two
livenessProbe Is it alive? Recreates the container
readinessProbe Can it take requests? Removes it from the endpoints

Until the startupProbe succeeds, the other two do not start at all. Without it, you have only two options.

A Job's two values

completions is how many times it must succeed, and parallelism is how many to run at the same time. If you mix them up, the Job never ends or uses more resources than necessary.

backoffLimit makes the retry interval grow exponentially, so with the default 6, it takes a few minutes to give up. Meanwhile, failed Pods keep piling up.

QoS is not something you set

You cannot write the qosClass field directly. The kubelet assigns it from the combination of requests and limits. And when the node runs short of memory, BestEffort is evicted first.

When configuration injection takes effect

Whether a Pod sees a change to a ConfigMap or Secret right away depends on how you injected it. A value injected as an environment variable is copied once when the container starts, so later changes are never reflected. You have to recreate the Pod. A file mounted as a volume, on the other hand, is refreshed periodically by the kubelet, so it changes to the new content after a short while, but it is useless if the application does not reread the file.

So there are two methods commonly used in practice. One is to put a hash of the configuration contents into an annotation in the Pod template, so that when the configuration changes the template changes too and a rolling update happens by itself. The other is to have the application detect file changes and reread them. The former is much simpler and easier to roll back, so it is better in most cases.

One more thing to know is the difference when you reference something that does not exist. If you reference a nonexistent ConfigMap through an environment variable, the container cannot be created and you get CreateContainerConfigError, but if you reference it as a volume, the mount never finishes and the Pod gets stuck in ContainerCreating. It is the same typo with different symptoms, so knowing this correspondence greatly shortens investigation time.

What to base probe values on

What is actually hard is not attaching a probe but which values to put in. It is not rare to kill your own service by using the defaults as they are, so you need to know what each value changes.

Value Meaning If set wrong
periodSeconds How often to ask Too short burdens the app; too long delays failure detection
timeoutSeconds How long to wait for a response The default is 1 second, so under load a healthy app is counted as failed
failureThreshold How many consecutive failures before acting At 1, even a transient delay causes a restart

The combination that causes the most incidents is a short liveness timeout. When load piles up and responses slow down, the probe fails, the container restarts, load piles onto the remaining Pods so they slow down too, and in the end everything restarts over and over. It amounts to trying to fix load-caused delay with restarts, so the situation worsens by itself. That is why the principle is liveness lenient, readiness sensitive. Being removed from traffic for a while can be undone, but a restart cannot.

What you check also matters. Liveness should look only at whether the process is in a state it cannot recover from by itself, so it must not check whether it can reach the database it depends on. When the database wobbles for a moment, all the application Pods restart, and then the connections all pile back in at once and make things worse. Dependency checks belong to readiness. If it fails there, only the traffic is removed, and when the database returns, the Pod quietly rejoins.

When you use a startupProbe for a slow-starting app, just remember that failureThreshold × periodSeconds is the allowed startup time. 20 tries every 30 seconds means it waits up to 10 minutes, and if it has not come up by then, restarts begin from that point.

What really matters in practice

For an app that starts slowly, the answer is a startupProbe. Without it, you have only two options — make liveness loose and leave a truly dead container alone for a long time, or make it tight and kill the app during startup so it restarts forever. The third probe was created to remove that dilemma.

The default backoffLimit of 6 is longer than you think. The retry interval grows exponentially, so it takes a few minutes to give up, and meanwhile failed Pods keep piling up. A Job that needs to fail fast must lower the value.

QoS is not something you write; it is something you are assigned. The kubelet decides it from the combination of requests and limits, and when the node runs short of memory, BestEffort is evicted first. Not writing requests for an important workload is the same as writing "you may kill me first."

In the next lab, you confirm all of these through behavior.