Break Kubernetes — My YAML Did It
Alive is not the same as ready
Summary
A readiness failure means losing eligibility to receive traffic; it does not mean the container process has exited.
Why this matters
Suppose a service takes time to load data or a cache when it starts. The process is alive, but it is not yet ready to handle normal requests. If you keep sending requests in this state, customers get errors, and if you restart unconditionally, the startup work may begin again from scratch. That is why you need to distinguish process liveness, completion of initial startup, and whether the service can accept traffic.
The Kubernetes liveness, startup, and readiness probes are the tools that express these distinctions. Use liveness to decide on restarts, startup to protect a slow startup period, and readiness to decide whether the service can receive traffic. Rather than memorizing the probe names, ask "what changes when it fails?" This lab has only a readiness probe, so explaining a failure as a restart caused by liveness would contradict what you observe.
How it works
The lab app returns a fixed body on path / at port 8080. The healthy readiness probe also checks the same /. If you change only the readiness path to /missing, the app itself still handles requests to /, but the probe receives an HTTP 404 from a path that does not exist. The Pod may be Running, yet the container's ready is false. Even if an address remains in the Service's EndpointSlice, a target whose ready condition is false is not a ready target for ordinary service traffic.
GET PodIP:8080/ → 200, 앱은 살아 있다
readiness /missing → 404, 수신 준비 조건은 실패한다
Pod phase → Running일 수 있다
container ready → false
Service의 Ready 대상 → 0
When these five lines appear together, "the probe contract differs from the real path, so the Pod dropped out of the traffic targets" is a better explanation than "the app died and the service stopped." However, you must not carry a 404 event from a past Pod over to the current Pod as its cause. Because the template change replaces the Pod, check the name and the UID together to confirm the event belongs to the current Pod. An old event is reference material that a failure occurred; it does not automatically prove the current cause.
A readiness failure alone does not restart the same container. This lab observes together the state where the current Pod's restartCount is 0 and a direct HTTP request succeeds. Do not confuse two things: changing the template's readiness path itself creates a new Pod, while a probe failure repeatedly restarting a container is a different matter. A new Pod UID is the result of the Deployment change, and restartCount inside that new Pod is the record of container restarts.
A typical Deployment uses RollingUpdate. If the new Pod does not become ready, the previous healthy Pod can remain and keep serving requests. This does not mean there is no failure; it means the rollout is blocked but the old version is protecting the service. This course uses a single replica and Recreate so that the old Pod does not hide the experiment's symptoms. It is a controlled condition for observing cause and effect, not a recommended setting for zero-downtime operation.
This is also why you leave every other setting unchanged and change one thing. If you break both the readiness path and the Service selector at once, there are two reasons the number of Ready targets is 0. Even if you recover only one, requests do not come back, and the learner may doubt the recovery command itself. Fully recover the selector experiment first, then move on to the readiness experiment. Compound failures are a task for after you can tell each single failure apart.
What it looks like in the field
In a new version, the health check path may change from /health to /ready while the deployment configuration still keeps the old path. The app's main business requests are fine, but only the new Pods never become Ready. Deleting the probe to make it green may not solve the cause. That is because you have removed the safeguard that judged whether the Pod could actually receive traffic. You need to align the path that the app and the deployment configuration agreed on and check that the state changes.
Conversely, there are cases where the probe honestly exposes a failure of a real dependency. If the database is unavailable and you change readiness to always return 200, the check passes but customer requests keep failing. A probe is a condensed version of the service contract. Which dependencies to include, how sensitive it should be to transient errors, and how costly the check is must be designed according to the service's characteristics. Copying the / from one lab to every service is not the answer.
If readiness and liveness call the same heavy business API, the load increases, and a vicious cycle can occur in which several Pods restart together even though a dependent service was only briefly slow. This course does not inject that kind of cascading failure into a production cluster. Still, you must learn that the two probes are responsible for different judgments before you can discuss thresholds and failure policies later. Shortening the check interval does not by itself make incident response always faster or safer either.
Practice writing sentences from your observations. Instead of "Kubernetes is broken," record "The current Pod is Running and a direct request to / succeeds, but that Pod's readiness /missing returns 404 and the number of Ready endpoints is 0." The latter sentence lets the next person make the same request and either refute or confirm it. It narrows the scope of the next check without rushing to assert a cause.
What to do in the next check
In the quiz you separate Pod replacement from container restart, and the purpose of a probe from its result. In steps 4–5 of the lab, you save the wrong readiness path and the recovery separately. The success of an observation tool means the experiment conditions were reproduced; it does not mean the tool fixed the probe for you. At the end, you connect the cause and the preventive measures in a report.