TT Lab
Get started
Learn Learning paths Courses

KCNA — Kubernetes and Cloud Native Associate

Running, but why is the service failing?

Continue in TT Lab

Goal

You distinguish Running, Ready, HTTP responses, and restarts through real kubelet behavior, and explain the illusions that arise when a check's meaning is wrong.

Why it matters

An open port does not guarantee business success. If you confuse an app that should stop accepting requests for a while with an app that should be restarted, automatic recovery makes the outage worse. While keeping a healthy comparison group, you observe readiness failure, process restart, the difference between TCP and HTTP checks, and slow initialization one by one. All experiments are limited to the synthetic app on a personal k3s VM. Do not bring in a production kubeconfig, real secrets, or external servers. This is a 55-minute lab, and preparation can take a few minutes. If you need more time, extend before the session expires. When it ends, the VM and files are reclaimed.

Prepared environment and helpers

In the namespace kcna-health, web-a and web-b are behind the healthy Service. The comparison lying Service initially has no targets. Both Services are ClusterIP:8080 with publishNotReadyAddresses=false. The app image runs with a pinned digest and Always pull, non-root, cap drop ALL, seccomp RuntimeDefault, a read-only root, and the ServiceAccount token not mounted. Only a 1Mi emptyDir at /state is used for the synthetic failure marker. Do not solve the task by changing workload security or the healthy comparison group.

The helper command is python3 /opt/fixtures/kcna_health_lab.py followed by act, capture, observe, grade, or prepare and a step number. act validates the input JSON and then makes only the designated experimental change. capture waits for the actual observation and saves it to /root/kcna-health. observe prints the current observation. grade changes neither resources nor student files. What the student creates is the change/probe input JSON. Do not fabricate the success, UID, or status values of the observation JSON. /opt/fixtures/kcna-health-context.json, kcna-health-lab-context.json, and kcna-health-warming.json are internal records managed by the helper.

Steps

  1. Save /root/kcna-health/baseline.json with capture 1. Investigate the Pod UID, containerID, restartCount, and Ready of web-a and web-b in kcna-health, the healthy Service and EndpointSlice conditions, and an actual HTTP 200. The Pods are non-root, cap drop ALL, read-only root, and token not mounted, and you preserve the healthy comparison group B to the end.
  2. Write a JSON with pod=web-a, flag=not-ready, present=true to /root/kcna-health/not-ready-change.json. Save not-ready.json with act 2 and capture 2. A must become Ready=false, direct HTTP 503, and EndpointSlice ready=false while staying Running, and the container and restart count must stay the same. Also look at whether six Service requests are delivered to the healthy B.
  3. Write pod=web-a, flag=not-ready, present=false to /root/kcna-health/restored-change.json. Save restored.json with act 3 and capture 3. Ready and healthy HTTP responses must return on the same Pod, container, and restart count.
  4. Write pod=web-a, flag=unhealthy, present=true to /root/kcna-health/restarted-change.json. Save restarted.json with act 4 and capture 4. On the same Pod UID, the container ID and the app's boot_id must change and the restart count must increase. This marker is a synthetic failure that is erased when a new process starts.
  5. Write tcpSocket.port=8080, periodSeconds=2, timeoutSeconds=1, failureThreshold=2, successThreshold=1 to /root/kcna-health/tcp-probe.json. act 5 creates the liar Pod with this configuration and puts in a readiness failure marker. Save tcp-lie.json with capture 5. TCP Ready=true and HTTP 503 of the lying Service must be observed at the same time.
  6. In /root/kcna-health/http-probe.json, use httpGet.path=/readyz and httpGet.port=8080 instead of tcpSocket, and keep the other four numbers the same as in step 5. act 6 confirms the original liar UID and replaces it with a new HTTP probe Pod. Save http-excluded.json with capture 6. Compare the direct HTTP 503, Ready=false, EndpointSlice ready=false, no healthy HTTP response from the Service, and the changed Pod UID.
  7. Write httpGet.path=/startupz, httpGet.port=8080, periodSeconds=1, timeoutSeconds=1, failureThreshold=45 to /root/kcna-health/startup-probe.json. act 7 creates the 20-second-initialization app slow and records started=false, Ready=false, restarts 0, and /livez 503 right after the start. Save warming.json with capture 7. Even if you press capture late, the actual initial observation is preserved.
  8. Save /root/kcna-health/started.json with capture 8. slow must be started=true, Ready=true, and HTTP 200 on the same Pod and container as in step 7 with no restart. The recovery of A, the liar excluded by the HTTP probe, and the original B's identity, container, and healthy responses must also be preserved.

Notes

If it fails, read observe and kubectl -n kcna-health get pods -o json on the same VM to investigate the cause. While waiting for the new state, do not call act repeatedly. Completed changes and observations are preserved, and past states are not recreated as current values. Initialization ends as time passes, so grading of step 7 looks at both the actual initial observation and the current same container with restarts 0. Step 8 additionally checks until the actual start completes. Ordinary grading keeps a 60-second budget and the whole step preparation a 90-second budget. Preparation runs only the earlier steps that are missing. It does not create the current inputs and observations for you, and it does not overwrite existing inputs you wrote incorrectly. The experiment's liveness marker is a synthetic failure erased on restart, and it is neither a real deadlock detector nor a recovery method for every failure. Do not use the standalone Pod replacement of the liar as is for zero-downtime deployment of a production Deployment. The observation files are learning material, not a remote attestation or cheating-prevention device that controls root. Official sources: Probe roles · EndpointSlice conditions · Probe configuration.

Investigate the baseline of the green light

Save /root/kcna-health/baseline.json with capture 1. Investigate the Pod UID, containerID, restartCount, and Ready of web-a and web-b in kcna-health, the healthy Service and EndpointSlice conditions, and an actual HTTP 200. The Pods are non-root, cap drop ALL, read-only root, and token not mounted, and you preserve the healthy comparison group B to the end.

Do not look only at the Running label; read the UID, container ID, conditions, and response each.

Exclude a running app from request assignment

Write a JSON with pod=web-a, flag=not-ready, present=true to /root/kcna-health/not-ready-change.json. Save not-ready.json with act 2 and capture 2. A must become Ready=false, direct HTTP 503, and EndpointSlice ready=false while staying Running, and the container and restart count must stay the same. Also look at whether six Service requests are delivered to the healthy B.

A readiness failure is not a termination command. Even if the address remains in the EndpointSlice, the ready condition can be false.

Restore the ready state without a restart

Write pod=web-a, flag=not-ready, present=false to /root/kcna-health/restored-change.json. Save restored.json with act 3 and capture 3. Ready and healthy HTTP responses must return on the same Pod, container, and restart count.

This is not a recovery that creates a new Pod. Clear only the readiness marker and keep the original process.

A container restart inside the same Pod

Write pod=web-a, flag=unhealthy, present=true to /root/kcna-health/restarted-change.json. Save restarted.json with act 4 and capture 4. On the same Pod UID, the container ID and the app's boot_id must change and the restart count must increase. This marker is a synthetic failure that is erased when a new process starts.

Rather than the fact that the Pod name is the same, compare how the Pod UID and container ID changed.

The port is open but HTTP is 503

Write tcpSocket.port=8080, periodSeconds=2, timeoutSeconds=1, failureThreshold=2, successThreshold=1 to /root/kcna-health/tcp-probe.json. act 5 creates the liar Pod with this configuration and puts in a readiness failure marker. Save tcp-lie.json with capture 5. TCP Ready=true and HTTP 503 of the lying Service must be observed at the same time.

Accepting a connection and a business response are different questions. An HTTP 503 is also delivered over a TCP connection.

An HTTP probe that checks business readiness

In /root/kcna-health/http-probe.json, use httpGet.path=/readyz and httpGet.port=8080 instead of tcpSocket, and keep the other four numbers the same as in step 5. act 6 confirms the original liar UID and replaces it with a new HTTP probe Pod. Save http-excluded.json with capture 6. Compare the direct HTTP 503, Ready=false, EndpointSlice ready=false, no healthy HTTP response from the Service, and the changed Pod UID.

Do not ignore the probe or expose addresses that are not ready. Replacing the probe of a standalone Pod requires a new Pod.

Protect and record slow initialization

Write httpGet.path=/startupz, httpGet.port=8080, periodSeconds=1, timeoutSeconds=1, failureThreshold=45 to /root/kcna-health/startup-probe.json. act 7 creates the 20-second-initialization app slow and records started=false, Ready=false, restarts 0, and /livez 503 right after the start. Save warming.json with capture 7. Even if you press capture late, the actual initial observation is preserved.

Do not mix the startup and liveness budgets. The observation helper keeps the initial state, so you do not race on click speed.

Prove the start completed without a restart

Save /root/kcna-health/started.json with capture 8. slow must be started=true, Ready=true, and HTTP 200 on the same Pod and container as in step 7 with no restart. The recovery of A, the liar excluded by the HTTP probe, and the original B's identity, container, and healthy responses must also be preserved.

Compare not only the current Ready but also the container ID and restart count of the initial observation.