Break Kubernetes — My YAML Did It
Recover three incidents and report the evidence
Summary
A successful rollback command is not the end of recovery. You have to confirm the healthy version you are going back to and send real requests again.
Why this matters
A small change to a container command made the app exit as soon as it started. The new Pod keeps restarting, and the on-call engineer wants to go back to the previous version. But was the "previous version" really healthy? If the previous experiment produced a template with a wrong readiness probe, simply stepping back one revision can be the road back to yet another failure. Set the recovery target not by memory or by sequence number but by a state you have confirmed to be healthy.
In this lab, after you confirm the readiness recovery, you record that Deployment revision. Next you change the command to a short shell that exits with the deliberate exit code 17. The 17 is a marker for this teaching scenario, not a common error code defined by Kubernetes. If the exit code differs, or if you see ImagePullBackOff, the cause is different from the failure intended this time, so do not treat it as the same answer.
How it works
A Deployment revision is updated when the Pod template changes and a rollout occurs. Changing a Service selector is not a change to the Deployment template. Changing only the replica count is also different from creating a new Pod template. So if you make up and memorize a rule such as "each kubectl command increases the revision by one," you will be wrong. Look at the rollout history and the contents of each revision to see what changed.
정상 템플릿 확인 → 정상 revision 기록
명령 변경 → 프로세스 종료 → 재시작과 backoff
증거 보존 → 원인·복구 목표 선택
정상 revision 지정 rollback
현재 Pod·Ready Endpoint·실제 HTTP 다시 확인
CrashLoopBackOff is a clue that indicates the state of waiting to restart. The word alone does not tell you whether it is an OOM, a wrong command, or a missing configuration file. Read the exitCode and reason in lastState.terminated, the restartCount, and the previous container's logs together. This command writes intentional-exit-17 to stderr and exits with 17. Only when the previous run's log contains that marker and the exit code also matches have you reproduced this experiment exactly.
Do not conclude that the app printed nothing just because the current container's log is empty. Right after a restart, the log you care about may be in the previous run. kubectl logs --previous is used to check this difference. However, logs are not a permanent audit store. They can disappear because of Pod replacement and log retention policies, so you save the parts you need into the observation file before recovering.
If you know the revision to roll back to, you can specify it with --to-revision. Even if you go back to a healthy Pod template, the Service is a separate object. If the Service targetPort is still 9999, requests to the Service can keep failing even after the Pod returns to Ready. That is why the final grading does not inspect only the report file or the Deployment state. It sends requests twice to the actual Pod IP and Service IP and checks for the expected response.
The request that confirms a successful recovery also needs to be clear about what it verified. In this course, GET / is a small HTTP response that uses no database or external payment. In a production recovery decision, you may need to check representative business paths, error rate, latency, dependencies, and data consistency together. Our successful response is evidence that this small service's connection path has recovered; it does not guarantee that all customer work is back to normal.
What it looks like in the field
An incident report is not a log of command execution. It has to connect what impact there was, which layer you suspected and how you narrowed it down, what you changed, and which request confirmed the recovery. A note such as "restarting the Pod fixed it" makes the next incident repeat the same action, but if you have evidence that it was a selector mismatch, you can build a preventive check that compares the selector and labels before deployment.
This report is JSON that connects cause, prevention, and evidence for the three incidents: selector, readiness, and crash. A machine does not score the style of free prose; it explicitly checks the cause distinctions and evidence links you are supposed to learn. The prevention codes are compare-selector-labels for service-selector, probe-real-endpoint for readiness-path, and smoke-test-command for process-command. The instructions give the meaning of each code, so this is not a task of guessing strings.
That said, a correct code in the report is not strong proof that you ran the experiment. The learner is root on their own VM and can edit the observation JSON. The saved observations are an educational experiment notebook and not a tamper-proof certificate. The current HTTP check in the last step is there at least to reject the case where someone dresses up only the files as healthy without fixing the service. Explaining the trust boundary of an exam honestly is also part of the quality of an operations document.
Because the same service moves between healthy, broken, and recovered over time, having every step look only at the current state causes a problem. When you click full grading after recovery, the earlier steps that told you to "create the broken state" would fail. This course preserves the observations from the first seven steps as separate files, and the last step checks their connections together with the current recovery. The observe command is a recording tool the student runs, and the grader neither edits those files nor injects failures.
Preventive measures also have to be confirmed by tests. If you deliberately put in a Service with the wrong port and the healthy verdict still passes, the check protects nothing. The fact that one healthy example passes is also true of a check that "always passes." The author has to confirm not only the correct answer but also that the check rejects a state where only the selector is fixed and the probe is wrong, a state with a different exit code, and a state where the Pod is Ready but Service HTTP fails.
What to do in the next lab
Over the 8 steps you observe the baseline and three kinds of failure and recovery, and at the end you tie cause, prevention, and evidence together. After the final success, run full grading again to confirm that the past evidence is preserved. The lab takes about 50 minutes and runs on a temporary VM. If you need more time, extend it before the session expires, and save important observations separately before the session ends. Do not copy this course's experiments as they are onto a production cluster.