KCSA — Kubernetes Security Associate
The Mesh's Identity, the Cluster's Certificates, Two Kinds of Eyes
In one line
A service mesh's mTLS gives each workload a service account-based identity and encrypts the wire, but it does not do authorization, and it cannot protect traffic the sidecar does not intercept. The cluster PKI has server and client certificates under three CAs (kubernetes-ca, etcd-ca, and front-proxy-ca), and client certificates created by kubeadm expire after one year. The audit log tells you "who asked the API for what," while runtime detection such as Falco tells you "what actually happened inside the node and containers," and the two fill in each other's blind spots.
Why this was needed
KCSA's platform security domain asks not about the cluster itself but about the layers placed on top of it. The claims that installing a mesh makes you safe, that kubeadm handles certificates on its own, and that turning on the audit log means you see everything — all three are only half right. If you do not know where they are right and where they are wrong, holes remain in the places where you believe a security control exists. The API server bypass risks document seen in the earlier module, stating that direct access to the kubelet API and etcd is not recorded in the audit log, is one example of such a hole.
How it works
The service mesh — what it does for you and what it does not
According to the Istio security concepts, the mesh tunnels service-to-service communication through client-side and server-side Envoy proxies. As the client Envoy performs an mTLS handshake with the server Envoy, it uses a secure naming check to confirm that the service account in the server certificate is authorized to run that service, and once the connection is up, the server Envoy authorizes the request. So what the mesh does for you is clear — the work of issuing, distributing, and rotating TLS certificates for each application, and confirming "which workload did this request come from" by service account identity rather than by IP.
The Istio security best practices document lists what it does not do. First, the default PERMISSIVE mode also accepts plaintext, so encryption is not guaranteed until you move to STRICT. Second, mTLS provides only authentication, so anyone with a valid certificate can reach the service, and to lock it down you must put AuthorizationPolicy in a default-deny pattern (for example, a conditional allow after an allow-nothing policy that denies everything). For a workload with no policy at all, Istio allows every request. Third, the sidecar intercepts only TCP, while UDP and ICMP pass through, some ports including 22 are excluded from inbound capture, and since the application and sidecar are in the same network and process namespaces, the application can delete the redirection rules and bypass the sidecar. The document's conclusion is "it is not safe to believe that all traffic is unconditionally captured," and so it recommends defense in depth by layering NetworkPolicy. AuthorizationPolicy selects its targets with selector or targetRefs, and if placed in the root namespace it applies to the entire mesh.
| What the mesh gives | What the mesh does not give |
|---|---|
| Workload identity (service account certificates) | Authorization — AuthorizationPolicy is needed separately |
| Encryption between sidecars | Protection between the application and the sidecar (plaintext inside the same Pod) |
| Rejecting plaintext in STRICT | Protection of UDP, ICMP, excluded ports, and bypassed traffic |
| Policy based on L7 attributes | The L3/L4 boundary — reinforced with NetworkPolicy |
PKI — three CAs and an expiry clock
The PKI certificates and requirements document divides the certificates a cluster needs into two groups. Server certificates are on the API server endpoint, the etcd servers, the kubelet on each node, and optionally the front-proxy. Client certificates are for when the kubelet authenticates to the API server, when the API server authenticates to etcd, when the controller manager, scheduler, and kube-proxy authenticate to the API server, and for administrators. There are three CAs that sign these.
| CA | File | Purpose |
|---|---|---|
| kubernetes-ca | /etc/kubernetes/pki/ca.crt |
The general Kubernetes CA |
| etcd-ca | /etc/kubernetes/pki/etcd/ca.crt |
All etcd-related certificates |
| kubernetes-front-proxy-ca | /etc/kubernetes/pki/front-proxy-ca.crt |
For the front-end proxy |
In addition, there is the key pair sa.key and sa.pub for signing service account tokens. To make the CAs hierarchical, you can create intermediate CAs from a single administrator-controlled root CA and leave the rest of the issuance to Kubernetes. What matters in the threat model is the weight of each CA — a client certificate signed by etcd-ca is access to all of etcd's data (the bypass risks document: "any certificate issued by a CA that etcd trusts allows full access to the data in etcd"), and if you create a certificate with O=system:masters with kubernetes-ca, it is a superuser. The document says the kube-apiserver's kubelet client certificate can use a less privileged group instead of system:masters, and that kubeadm uses the kubeadm:cluster-admins group.
SAN is also a security item. If you connect with a name that is not in the certificate's hosts (SAN), TLS verification fails, so the kube-apiserver certificate includes the hostname, host IP, and advertise IP, plus the load balancer's address and kubernetes, kubernetes.default, kubernetes.default.svc, kubernetes.default.svc.cluster, and kubernetes.default.svc.cluster.local. A typical incident is that when you later change the load balancer address, connections break because the SAN is missing.
The first sentence of the kubeadm certificate management document is the clock — client certificates created by kubeadm expire after one year. kubeadm certs check-expiration shows the expiry of the certificates in /etc/kubernetes/pki and the client certificates embedded in admin.conf, controller-manager.conf, and scheduler.conf, and in the document's example output, the CA shows 9 years remaining. There are two ways to renew. kubeadm renews all certificates during a control plane upgrade, so if you upgrade once within a year there is nothing else to do (to turn it off, --certificate-renewal=false), and manually you use kubeadm certs renew, which for a replicated control plane must be run on every node and the control plane Pods must be restarted after running it (dynamic reloading does not happen). The same document also notes that the kubelet's serving certificate is self-signed by default, so external services cannot verify it with TLS.
Security observability — audit logs and runtime detection
According to the auditing document, auditing records the cluster's activity in chronological order so that it can answer what happened, when, who did it, to what, and where. The record begins inside the kube-apiserver. Events are generated at each stage of a request (RequestReceived, ResponseStarted, ResponseComplete, and Panic), the policy decides whether to record and at what level, and a backend (a log file or a webhook) stores it. The policy compares rules in order, and the first rule that matches decides the level, which is one of four: None, Metadata, Request, and RequestResponse. If you do not give --audit-policy-file, nothing is recorded, and a policy with zero rules is illegal. Auditing increases the API server's memory use.
The audit log's blind spots are stated exactly in the bypass risks document — direct access to the kubelet API, direct access to etcd, and the runtime socket pass through neither admission nor the audit log. The audit log shows only what the API server saw.
The eye that sees those blind spots is runtime detection. According to the Falco documentation, Falco is a CNCF graduated project providing runtime security in host, container, Kubernetes, and cloud environments, which parses the kernel's system calls in real time, checks them against a rules engine, and raises an alert when a rule is violated. It attaches container runtime and Kubernetes metadata to events to tell you "in which Pod of which namespace," and can also take event sources outside system calls through plugins. As what the default rules catch, the document lists privilege escalation using privileged containers, namespace changes using tools like setns, and reads and writes to well-known directories. Alerts are sent to a SIEM or data lake for use in investigation.
The two eyes see the same event from different places. When someone execs into a Pod, the audit log records the request to the pods/exec subresource and the requester, and Falco records the system call that a shell appeared inside that container. If a shell appears only in Falco and not in the audit log, you can suspect a path that did not go through the API server (the kubelet API, the runtime socket, or a static Pod), and if it is in the audit log but Falco has nothing, the request was rejected or has not run yet. If you turn on only one, this comparison is impossible.
What it looks like in the field
One day, a year later, kubectl gives an x509 error. This is the case where a kubeadm cluster was set up and went over a year without an upgrade. The certificate in admin.conf and the certificates of the control plane components expire on the same day, so even if the API server itself is running, nobody can talk to it. If you put check-expiration in regular checks, you can see the remaining time months in advance.
A compromise with no trace in the audit log. If an attacker who SSHed into a node started a container through the runtime socket, the API server sees nothing. In this case the only records are runtime detection watching the node's system calls and the node's authentication logs, which is why the bypass risks document says to restrict node access itself.
What to check in the next quiz
The quiz asks about what mTLS does not give, traffic a sidecar cannot intercept, the roles of the three cluster CAs, the expiry and renewal of kubeadm certificates, how an audit policy evaluates rules, and what the audit log and runtime detection each see.