TT Lab
Get started
Learn Learning paths Courses

KCNA — Kubernetes and Cloud Native Associate

What green means and how to budget for slow startup

Continue in TT Lab

One-line summary

A probe's success means that the chosen check condition is true. If you interpret "the port is open" as business success, or interpret slow initialization as a failure that must be killed immediately, automatic recovery creates an outage instead.

Why this was needed

A restaurant whose phone connects is different from a restaurant that can take orders. The staff answer the phone, but the kitchen may have stopped. The moment you read "connection succeeded" as "order succeeded" in monitoring, a green screen and customer complaints exist at the same time. A server's health endpoint is not automatically a correct check just because the word health is in its name. You have to look at who implemented the success condition and with what.

How it works

A TCP check verifies whether a connection to the specified port can be made. An HTTP check observes the response code of the specified path. For example, even if a web server returns 503 for every order, the socket that accepts TCP connections may be open. In this case the TCP readiness succeeds while users receive failure responses.

That does not mean the TCP check is wrong for every app. If the purpose of the check is whether connections are accepted, it is the exact question. The problem is attaching business meaning that the result does not guarantee. The same goes for an HTTP check: a path that always returns 200 will not notice a database query failure. Conversely, if every health check performs a heavy query, the check itself can burden the service.

Example path What it tries to confirm What should not be mixed in carelessly
/startupz Has this process finished its initialization? Repeating the same initialization every time after it has already started
/livez Does it make sense to keep it running? Temporary failures of all external dependent services
/readyz Can it be handed new business requests? A guarantee of complete success of the entire user journey
Actual business path The actual request result and response content Extending a single success into an availability guarantee

The startup probe is a separate check that protects slow starts. Until it succeeds, liveness and readiness do not run. If you apply an aggressive liveness from the start to an app whose initialization takes time, the app can keep dying before it even finishes starting. This is why the time allowed for starting and the time for detecting an unrecoverable state during running are handled separately.

For example, an experimental app gives an initializing response for 20 seconds after it starts, and then gives normal responses. If the startup check period is 1 second and the consecutive-failure limit is 45, that setting leaves margin over the 20-second initialization. These numbers are not a guarantee that it is terminated at exactly 45.000 seconds in every environment. There are execution, scheduling, and check overheads, so you have to observe the actual transition. In an exam too, being able to explain which budget protects which failure matters more than memorizing the numbers.

A brief readiness failure after initialization has finished does not make startup run again from the beginning. Conversely, when a container is restarted, a new process has a start phase. You must not reuse the previous process's success record as evidence for the new process. That is why we record not only the Pod UID but also the container ID and the app's run identifier.

What it looks like in practice

Suppose that after a new version is deployed, the connection check succeeds but only the real API fails. First, you do not remove the probe or change the Service to expose addresses that are not ready. You investigate separately the difference between the business path and the check path, the Pod conditions, the EndpointSlice targets, and the actual request results. Then you fix the check so that it expresses the readiness that the service promises. The red light disappearing from the check is different from the outage being resolved.

When you change a check, you also have to respect object lifetimes. You cannot freely modify the probe fields of a running standalone Pod. You modify the declaration and confirm which managing object will create the new Pod and whether there is an available comparison group during the replacement. Do not use the textbook's disposable single-Pod replacement procedure as is for zero-downtime deployment of a production Deployment.

What you will do in the next lab

You fail the readiness of one of two healthy Pods and see that only the request assignment changes, without a restart. Next you observe a container restart through a synthetic liveness failure, and compare a separate Pod that passes the TCP readiness while returning HTTP 503. Finally, you connect the state during the slow start and the state after initialization completes to actual records. A transition can take time, so before you overwrite a setting repeatedly, look at the progress of the same object.

Official source: Probe configuration and caveats.