Break Kubernetes — My YAML Did It
My YAML did it — a casebook of three failures
Goal
On a real k3s cluster, you create a selector mismatch, a wrong readiness path, and a process exit, save the observation evidence, and then recover with minimal changes. This is a 50-minute experiment for intermediate learners who know kubectl get/apply and the basic structure of YAML.
Why it matters
Running, Ready, and a user's request succeeding are different facts. Even with a healthy Pod, customers fail if the Service selector or port is wrong. Conversely, a process can be alive yet fail its readiness condition and be excluded from the forwarding targets. If you change one field at a time and compare before and after, you can explain the cause instead of restarting over and over. The scope of this lab is to tell the three incidents apart and then return to a revision you have confirmed to be healthy.
Safety scope
Inside the temporary VM, use only the API at https://127.0.0.1:6443과 the labhub-mystery namespace. Do not run this on a production cluster or in another student's environment. The provided app uses a single replica and Recreate so that the previous healthy Pod does not hide the failure. This does not mean it is a recommended setting for zero-downtime operation. The app is pinned by image digest. When the session ends, the VM and the observation files disappear. If needed, extend the time before it expires and save the observations separately. The observation files are editable experiment notes, not tamper-proof evidence.
Steps
- Apply the provided /opt/fixtures/k8s-mystery-app.json with kubectl apply. Query the parcel Deployment and Service in the labhub-mystery namespace, and record as baseline the healthy state in which the / path on port 8080 of both the Pod and the Service returns the body labhub-mystery-v1. Recording command:
python3 /opt/fixtures/k8s_mystery.py observe baseline /root/k8s-mystery/01-baseline.json. - Change only the selector of the parcel Service to app=missing. Keep the Pod label and the Deployment unchanged. Record as selector-broken the state in which the Pod is Ready and responds directly, but the Service has 0 selected endpoints and requests fail. Recording command:
python3 /opt/fixtures/k8s_mystery.py observe selector-broken /root/k8s-mystery/02-selector.json. - Restore the Service selector to app=parcel. Confirm 1 selected Ready endpoint and healthy HTTP on both the Pod and the Service, and record it as selector-restored. Do not delete and recreate the Deployment. Recording command:
python3 /opt/fixtures/k8s_mystery.py observe selector-restored /root/k8s-mystery/03-selector-fixed.json. - Change only readinessProbe.httpGet.path of the Deployment's parcel container to /missing. Record as readiness-broken the state in which the new Pod is Running and responds directly on /, but has Ready=false, 0 restarts, a 404 event on the current Pod, and 0 Ready endpoints. Recording command:
python3 /opt/fixtures/k8s_mystery.py observe readiness-broken /root/k8s-mystery/04-readiness.json. - Restore the same readiness path to /. Do not delete the probe or replace it with an exec that always succeeds. Record the healthy HTTP on the Pod and Service and 1 Ready endpoint as readiness-restored. Recording command:
python3 /opt/fixtures/k8s_mystery.py observe readiness-restored /root/k8s-mystery/05-readiness-fixed.json. - Look up the healthy Deployment revision and save it in the Deployment annotation labhub.io/known-good-revision. Then replace the parcel command with ["sh","-ec","echo intentional-exit-17 >&2; exit 17"]. Confirm CrashLoopBackOff, the previous exit code 17, a restart count of 1 or more, and intentional-exit-17 in the previous log, and record it as crash. Recording command:
python3 /opt/fixtures/k8s_mystery.py observe crash /root/k8s-mystery/06-crash.json. - Check the rollout history and the saved labhub.io/known-good-revision, then run rollout undo to that revision. Keeping the existing Deployment, record as rollback the state in which the healthy command, the probe on /, the Service selector app=parcel, targetPort 8080, and HTTP on both sides are recovered. Recording command:
python3 /opt/fixtures/k8s_mystery.py observe rollback /root/k8s-mystery/07-rollback.json. - Keep the first seven observation files and, in /root/k8s-mystery/report.json, link cause, prevention, and evidence for each of the selector, readiness, and crash incidents. The cause and prevention codes are in the reference below, and evidence is the name of the file from the time of failure. The current Service must also be healthy. At the end, click full grading to check the past observations and the current recovery together.
Reference
The observe command of the observation tool only queries and saves files. It does not inject the requested failures or perform recovery for you. If you do not see the condition within 45 seconds, check the values and events and observe again. Grading of the first seven steps reads the saved records, so the earlier steps are not cancelled even after recovery. The final grading checks the record links of the same Deployment and the current Pod and Service HTTP twice.
- Selector incident: service-selector / compare-selector-labels — compare the selector against the Pod labels in advance.
- Probe incident: readiness-path / probe-real-endpoint — check the health-check path that is actually served.
- Exit incident: process-command / smoke-test-command — a small execution test of the container start command.
- Use
kubectl -n labhub-mystery get pods,svc,endpointslices -o wideto look at the targets and their states separately. - Link the events in
kubectl -n labhub-mystery describe pod <이름>(where the placeholder is the Pod name) to the current Pod UID. kubectl -n labhub-mystery logs <이름> --previous(where the placeholder is the Pod name) shows the previous container's log.- Deleting the probe, recreating the Deployment, and rolling back to a memorized number all avoid the contract of the problem.
The baseline where parcels arrive normally
Apply the provided /opt/fixtures/k8s-mystery-app.json with kubectl apply. Query the parcel Deployment and Service in the labhub-mystery namespace, and record as baseline the healthy state in which the / path on port 8080 of both the Pod and the Service returns the body labhub-mystery-v1. Recording command: python3 /opt/fixtures/k8s_mystery.py observe baseline /root/k8s-mystery/01-baseline.json.
Do not look only at the Running indicator; distinguish Ready, EndpointSlice, and actual HTTP. The observation tool does not create the state you want.
The delivery truck is alive but has no destination
Change only the selector of the parcel Service to app=missing. Keep the Pod label and the Deployment unchanged. Record as selector-broken the state in which the Pod is Ready and responds directly, but the Service has 0 selected endpoints and requests fail. Recording command: python3 /opt/fixtures/k8s_mystery.py observe selector-broken /root/k8s-mystery/02-selector.json.
Compare the Service's selector with the Pod's labels. Changing the selection condition does not change the existing Pod labels along with it.
Fix only the delivery list
Restore the Service selector to app=parcel. Confirm 1 selected Ready endpoint and healthy HTTP on both the Pod and the Service, and record it as selector-restored. Do not delete and recreate the Deployment. Recording command: python3 /opt/fixtures/k8s_mystery.py observe selector-restored /root/k8s-mystery/03-selector-fixed.json.
Do you have any grounds to restart the app? Restore only the field of the object you changed wrongly, and confirm the effect with a real request.
Deliveries halted because of a nonexistent health clinic
Change only readinessProbe.httpGet.path of the Deployment's parcel container to /missing. Record as readiness-broken the state in which the new Pod is Running and responds directly on /, but has Ready=false, 0 restarts, a 404 event on the current Pod, and 0 Ready endpoints. Recording command: python3 /opt/fixtures/k8s_mystery.py observe readiness-broken /root/k8s-mystery/04-readiness.json.
You change the readiness path, not liveness. Also check whether the event's target UID equals the current Pod's.
Match the health-check contract to the real path
Restore the same readiness path to /. Do not delete the probe or replace it with an exec that always succeeds. Record the healthy HTTP on the Pod and Service and 1 Ready endpoint as readiness-restored. Recording command: python3 /opt/fixtures/k8s_mystery.py observe readiness-restored /root/k8s-mystery/05-readiness-fixed.json.
Make the real request and the probe express the same health state. This case is a path typo, not a case of ignoring a real dependency failure.
The courier leaves a note with 17 and clocks out
Look up the healthy Deployment revision and save it in the Deployment annotation labhub.io/known-good-revision. Then replace the parcel command with ["sh","-ec","echo intentional-exit-17 >&2; exit 17"]. Confirm CrashLoopBackOff, the previous exit code 17, a restart count of 1 or more, and intentional-exit-17 in the previous log, and record it as crash. Recording command: python3 /opt/fixtures/k8s_mystery.py observe crash /root/k8s-mystery/06-crash.json.
Do not memorize the revision as a number like 1. Save it right after confirming the state is healthy, and read the evidence for the previous exit with logs --previous.
Return to the revision you confirmed, not the one you remember
Check the rollout history and the saved labhub.io/known-good-revision, then run rollout undo to that revision. Keeping the existing Deployment, record as rollback the state in which the healthy command, the probe on /, the Service selector app=parcel, targetPort 8080, and HTTP on both sides are recovered. Recording command: python3 /opt/fixtures/k8s_mystery.py observe rollback /root/k8s-mystery/07-rollback.json.
A Deployment rollback does not fix the Service. Check the number in the history together with a real request.
Close the case on evidence, not on a green light
Keep the first seven observation files and, in /root/k8s-mystery/report.json, link cause, prevention, and evidence for each of the selector, readiness, and crash incidents. The cause and prevention codes are in the reference below, and evidence is the name of the file from the time of failure. The current Service must also be healthy. At the end, click full grading to check the past observations and the current recovery together.
Keep the selector, the probe path, and the process command distinct from one another. If the current service fails, changing only the JSON will not pass.