TT Lab
Get started
Learn Learning paths Courses

KCNA — Kubernetes and Cloud Native Associate

Running, Ready, and business success are different questions

Continue in TT Lab

One-line summary

Running is a summary of the Pod lifecycle, and Ready is a condition about whether it is OK to hand requests to it right now. You have to look at the restart count and the actual HTTP response together to tell "the process is alive" from "the service is doing its work."

Why this was needed

A warehouse employee having shown up for work does not mean delivery preparation is finished. The employee is alive and moving, but the label printer may not be connected. The question of whether someone showed up is different from the question of whether to assign a new order. A server, too, can read its configuration or recover its connections while its process is up.

If an operator looks only at Running on the screen and judges that there is no outage, the users' failures go unexplained. Conversely, if you keep killing a process because one response was late, requests pile up on instances that are still healthy, and those instances slow down in turn. Health signals are not a feature for producing many green lights; they are an agreement about which action to attach to an observation result.

How it works

A Pod has no single "health score." status.phase, status.conditions, and per-container states answer different questions. The short table of kubectl get pods is a summary, so you must first check which column you are looking at. The fact that a restart occurred and it is Running now does not mean the past failure has gone away.

The Running phase also includes the case where the Pod has been assigned to a node and its containers have been created, and at least one is running or is starting or restarting. It does not mean every container's work is normal. Also, the CrashLoopBackOff you see in the table's STATUS and the API's status.phase are not the same field. Rather than memorizing a single summary string as a health verdict, unfold and read down to the per-container state.

Observation The question it answers What it cannot tell you alone
phase=Running Has the Pod entered the running phase? Whether business requests succeed
Ready condition Is it ready to receive traffic now? Whether every user scenario succeeds
restartCount How many times has this container been restarted? Which failure was the cause of the restart
Actual HTTP response Did this request on this path succeed? Whether other paths and other times succeed too

The readiness check is performed by kubelet, and its result is reflected in the Pod conditions and in the Service's delivery targets. In an ordinary Service, a Pod that is not ready is not used as a valid target for new requests. Here, do not look only at whether the address literally disappeared from the EndpointSlice. The address may remain with conditions.ready=false. Only by checking the Pod UID, the address, and the condition together are you looking at the same target. A Service with publishNotReadyAddresses=true behaves differently, so this lab explicitly sets it to false.

Suppose Pods A and B are behind the same Service. If you fail only A's readiness, A may still be Running. You can compare how a request sent directly to A connects but fails at the business level, while new requests through the Service are handled by B. The important evidence here is not just "I saw a response from B." You also have to check together that A's Ready=false, A's condition in the EndpointSlice, and A's original container ID and restart count are preserved.

Liveness, by contrast, is a check about whether to keep the container alive. When the configured consecutive-failure threshold is exceeded, kubelet terminates that container and follows the restart policy. When a container comes back up inside the same Pod, the Pod UID is the same and the container ID differs. This is a different event from a Deployment replacing the Pod. If you miss this distinction, you reach the wrong conclusion "the Pod is unchanged, so there was no restart."

What it looks like in practice

The first situation is a temporary failure of a dependent service. If the application itself is alive and can retry its connections, pausing request assignment for a while may be appropriate. When the dependent service recovers, the same process becomes Ready again. You cannot fix the other service's failure by killing your process.

The second situation is an application that cannot make progress internally. If it makes no progress at all even when it accepts new requests, and it is in a state recoverable by restarting the process, liveness can help. However, the failure marker in this lab is not a real deadlock but a synthetic failure for comparing recovery behavior. We do not claim from this result that we detected a deadlock in a production app. What a real app's check path should measure has to be designed separately.

Connecting to the next lesson

In the next reading, you look at the illusions that arise when the check itself is badly designed. You distinguish TCP connection success from HTTP business success, and slow initialization from unrecoverable failure. In the lab afterward, you record Pod and container identities, create a readiness failure and a liveness failure each, and then check whether the healthy comparison group was affected.

Official sources: Pod lifecycle and phase · Pod health checks · EndpointSlice conditions.