Break Kubernetes — My YAML Did It
Why did the delivery stop?
Summary
That a Service exists, that a Pod is running, and that a user's request succeeds are three separate pieces of evidence.
Why this matters
A parcel-tracking service has frozen. The on-call engineer runs kubectl get pods, sees Running, and reports that "Kubernetes is fine." Yet customer requests keep failing. Whose observation is wrong? Both can be true. Running is a state in the Pod lifecycle; it does not mean the address the user requested returned the right response. When you lump the answers to different questions into a single word, your incident analysis heads in the wrong direction.
In this course you become the culprit yourself. Inside a lab VM you build a small HTTP service, then change the Service selector, the readiness path, and the container command one at a time. You predict the symptom, observe the actual response, narrow down the cause, and recover. The only experiment target is parcel in the labhub-mystery namespace. None of the tasks stops the host cluster's nodes, CoreDNS, CNI, or etcd, and there is no reason to touch anyone else's server.
Start with a working knowledge of YAML indentation, the basic roles of Deployment and Service, and how to use kubectl get, describe, and patch. This is not a course for learning the command line from scratch; it is an intermediate lab that teaches you a command can succeed and the work can still fail. Every failure is created deliberately by the learner, and you record the evidence from the moment of failure before you start recovery.
How it works
A Deployment declares the Pod template you want and the number of replicas. A controller realizes that declaration by creating a ReplicaSet and Pods. A Service plays a different role. It picks target Pods with a selector such as app=parcel and provides a stable ClusterIP and port. The EndpointSlice shows the addresses the Service selected and their readiness. Creating a Service object does not automatically select the right Pods.
Service selector ── 비교 ── Pod metadata.labels
│ 일치한 대상
▼
EndpointSlice의 주소와 ready 조건
│ 대상 포트로 전달
▼
컨테이너의 HTTP 응답
In this lab, the normal values are app=parcel on both the Service and the Pod, and 8080 for both the service port and the target port. For a request to /, the app returns a single line, labhub-mystery-v1. The reason to check the response body instead of only whether the port is open is to tell apart a successful response from another process or from the wrong service. Even the number 200 is not enough evidence unless you know which request you made and which body you checked.
What changes first if you change the Service selector to app=missing? You did not change the Deployment's Pod template, so there is no reason for the Pods to restart. A direct request to the app's Pod IP still gets a response. But no Pod matches the selector, so the target disappears from the Service's EndpointSlice and requests to the Service IP fail. Rebuilding the container image or deleting the Pod at this point does not match the cause.
Be careful with programs that read an empty EndpointSlice, too. endpoints can be an empty array or can be represented as null. Treat both as "no selected addresses"; do not give up on observing altogether because of a parser exception. Conversely, do not assume everything is fine just because one address exists. As you will see in the next module, an address with ready set to false can be recorded. Separate the existence of an address from the number of targets that can actually be routed to.
During an experiment, a change does not show up everywhere at once. Right after you modify the Service, the controller needs a short time to update the EndpointSlice. An exit code of 0 from a command means the API accepted the change; it does not mean the whole data path has finished reflecting it. The observation tool rereads the state for a set period, but it neither fixes errors nor forges state. If the condition is still not met after that time, compare the value you changed with the value you observed again.
What it looks like in the field
If requests fail after a rollout, first narrow the scope. Check whether all requests fail or only a specific path, whether all Pods are affected or only some, whether requests also fail through the Service IP, and whether the Pod IP responds. If you guess at DNS, Ingress, TLS, and application dependencies all at once, you end up with too many hypotheses. This lab compares a ClusterIP request, which does not go through a DNS name or an external Ingress, with a direct request to the Pod. So passing here does not mean you have verified DNS and TLS on the internet.
For example, if the direct Pod response succeeds and only the Service fails, the hypothesis that the server process is completely dead becomes weaker. Next you look at the selector, the EndpointSlice, and targetPort. If the selectors match and a Ready address exists but only the Service fails, you can ask whether the target port number equals the port the process actually listens on. On the other hand, if the direct Pod request also fails, you have grounds to check the process's listen address, port, logs, and exit status first.
Evidence that supports a hypothesis is different from evidence that confirms it. A successful request to the Pod IP does not mean every network layer is healthy. Only the path of the request you observed succeeded. In production there may be additional replicas, NetworkPolicy, a service mesh, and an external load balancer. This course is a basic experiment that links symptoms to causes by changing one variable; it is not a universal diagnostic method that sorts every failure into three types.
The connection to a job role also needs a clear scope. The Canonical SRE job posting checked on 2026-09-10 covers Linux, Python, networking, and Kubernetes operations skills. The observe, isolate-the-cause, and recover exercises in this course translate those requirements into learning tasks; they are not a company's hiring test or a partner program. The focus is less on memorizing many commands and more on being able to explain which layer you inspected and to make your observation reproducible by the next person.
Compare the detailed diagnostic order with the official Kubernetes guide to debugging Services. The ports, response body, and namespace in this lab are values chosen for teaching.
What to do in the next check
In the quiz you distinguish the meanings of Running, EndpointSlice, and an actual HTTP response. In steps 1–3 of the last module's lab, you save the baseline, the selector error, and the selector recovery, each as observation JSON. Once you fix the failure you can no longer see the empty target list from that moment, so record it before recovering. That record is an experiment notebook the learner writes; it is not a tamper-proof audit log.