KCNA — Kubernetes and Cloud Native Associate
Observability — Three (or Four) Signals
One-line summary
Monitoring is "answering the questions you decided on in advance," and observability is "the property of being able to answer questions you did not think of in advance." The signals are separated because each one can answer different questions.
Why this was needed
In the single-server days, a single log file was enough. In microservices, one request passes through ten services, and logs alone cannot tell you which of them is slow. On top of that, Pods die and come up anew, so you start investigating after the one that left the logs is already gone.
So the approach of "when an outage happens, log into the server and look" does not hold. There is no server to log into in the first place, or it has already been replaced. You have to send the data out.
How it works
Three pillars, and a fourth
| Signal | The question it answers | Cost | Representative tools |
|---|---|---|---|
| Metrics | How bad is it now? Since when? | Cheap (aggregated numbers) | Prometheus |
| Logs | What exactly happened at that time? | Medium to expensive | Loki, Elasticsearch |
| Traces | Where did this request spend its time? | Expensive (usually sampled) | Jaeger, Tempo |
| Profiles | Who is eating this process's CPU/memory? | Expensive | Pyroscope |
The order of investigation is usually this order too. Detect the anomaly and narrow the scope with metrics, find which segment with traces, and read the story of that segment with logs. If you start digging through logs, you end up looking for a needle in a haystack.
The collection path in Kubernetes
- Metrics: each component and app exposes
/metrics, and Prometheus pulls periodically - Logs: the container's stdout/stderr → a file on the node's disk → a DaemonSet collector reads and ships it
- Traces: the app is instrumented and pushes over OTLP → a collector → a backend
One point KCNA stresses here: container logs must be sent to stdout/stderr. If an app writes to its own file, it disappears when the Pod dies and the collector cannot see it either. The 12-factor item "logs are an event stream" is exactly this.
The problem OpenTelemetry solves
In the past, each backend had a different SDK. If you used Jaeger and then switched to something else, you had to change the application code again. OTel cut this dependency by standardizing the instrumentation API/SDK and the transport protocol (OTLP). You instrument once, and you change the backend with collector configuration.
The observability material Kubernetes gives you from the start
- Events: The lower part of
kubectl describe pod. The control plane's grounds for its decisions, such as scheduling failures, image pull failures, and probe failures, are left here. It is the place to look before you dig through logs. - Status:
.status.conditionsand.status.phase. - metrics-server: The minimal resource usage that
kubectl topuses. It does not keep data long term.
Reliability terms
- SLI: The indicator you actually measure (for example, the fraction of successful requests)
- SLO: The target for that indicator (for example, 99.9% over 30 days)
- SLA: The contractual compensation when that target is missed. A contract concept, not a technical one
- Error budget: If the SLO is 99.9%, 0.1% is a budget you are allowed to break. When this budget remains, you deploy faster, and when it runs out, you focus on stabilization, so it becomes a basis for decisions
What it looks like in practice
There are two cases in the author's homelab where observability showed up dramatically.
First, Cilium's Hubble. It observes directly in the kernel with eBPF without injecting a single sidecar, and the actual output looks like this.
07:50:38.064: ebpf-demo/client:47918 -> ebpf-demo/api:80 http-request FORWARDED (HTTP/1.1 GET http://api/)
07:50:38.223: ebpf-demo/client:47932 -> ebpf-demo/api:80 http-request DROPPED (HTTP/1.1 POST http://api/)
07:50:38.223: ebpf-demo/client:47932 <- ebpf-demo/api:80 http-response FORWARDED (HTTP/1.1 403 0ms POST)
The method, path, response code, and latency are all visible, and DROPPED by policy is explicit. The point is that you can know "why was it blocked" from a record rather than a guess. The nginx in the same cluster was GET 200 / POST 200 before the policy was applied and split to GET 200 / POST 403 after, and that change remained in the flow log as it was.
Second, the KubeVirt case from earlier. The status of every component was AllComponentsReady, but the VM did not start. A status field says "I have been turned on," not "what I do actually works." So the author built the cluster verification not from status lookups but from 14 real-behavior scenarios (nodes Ready, control plane components, CNI, CoreDNS, worker scheduling, Pod-to-Pod communication, service DNS + HTTP 200) and confirmed 통과 14 / 실패 0 (14 passed, 0 failed). The final form of observability is the synthetic check.
What to check in the next quiz
With the final quiz of this course, you check the CNCF ecosystem and observability together. Next is KCSA, a course that looks again through an attacker's eyes at the structure you have learned so far.