CKS — Kubernetes Security Specialist
Encryption on the Wire, Vulnerabilities over Time, Surface on the Node
In one line
There are two ways to encrypt traffic between Pods — Cilium's transparent encryption wraps it in WireGuard or IPsec in the node kernel, and Istio's mTLS terminates in a sidecar inside the Pod. Running a version whose support has ended means you won't receive CVE fixes, so upgrading is itself a security item, and for the node OS the principle is to keep only what is needed to run the kubelet, runtime, and CNI.
Why this was needed
A NetworkPolicy decides only who can connect to whom. After a connection is allowed, the bytes that pass over the wire are plaintext. Within the same node everything is visible from the node, so encryption is meaningless, but between nodes traffic passes through physical switches, cloud virtual networks, and the equipment of other tenants. The security checklist says "not every CNI plugin provides encryption in transit, and if the plugin you chose lacks that feature, a service mesh is the alternative." That is why the CKS asks about the two routes, Cilium and Istio.
The reason upgrading is a security item is time. A vulnerability can be blocked only with a fixed version, and the project patches only the latest three minor versions. A cluster outside that window has no fix to receive even when a CVE is published. The node OS follows the same principle. Every installed service, package, and kernel module is a door an attacker can knock on, and removing doors you don't use is the cheapest defense.
How it works
Cilium transparent encryption — in the node kernel
According to the Cilium transparent encryption documentation, Cilium encrypts traffic between the endpoints it manages with IPsec or WireGuard (and, as a beta, with ztunnel). The application knows nothing. This is because encryption and decryption happen in the node's kernel datapath.
When you turn on WireGuard, each node's agent generates its own key pair and announces the public key through the network.cilium.io/wg-pub-key annotation on the CiliumNode object. The other nodes use that public key to set up a tunnel with that node. The tunnel endpoint is UDP 51871, so in environments with firewalls this port must be open between nodes, and in tunnel routing mode, packets wrapped once in VXLAN or Geneve are wrapped once more by WireGuard. The Helm values are encryption.enabled=true and encryption.type=wireguard, and to include traffic between nodes and between a node and Pods, you add encryption.nodeEncryption=true. The kernel must support WireGuard (it is built in from Linux 5.6 and later).
IPsec distributes keys as Kubernetes Secrets, and since ESP traffic flows between nodes, you must open ESP in the security group or firewall. From Cilium 1.18, in tunnel mode IPsec is applied after the overlay encapsulation, so even the security identity used for policy is encrypted on the wire. There are also limits the documentation states — it isn't supported in configurations chained on top of another CNI, and IPsec decryption is limited to one CPU core per tunnel.
The two approaches share two properties. First, packets going to the same node are not encrypted. The design reasoning is that the original is visible on the node, so it is meaningless. Second, whether to encrypt is decided by judging "is the destination a remote Cilium endpoint," the same way policy enforcement does, so in the brief moment before information about a new endpoint has spread, the first packet may go out in plaintext. The documentation suggests as countermeasures a policy that blocks egress to reserved:world, or encryption.strictMode.
Istio mTLS — in the sidecar inside the Pod
The flow in the Istio security concepts documentation is this. The client's request is redirected to the sidecar Envoy inside the Pod, and while the client Envoy does an mTLS handshake with the server Envoy, it uses a secure naming check to confirm that the service account in the server certificate is authorized to run that service. Once the connection is up, the server Envoy authorizes the request, and if it passes, it hands it to the backend over local TCP. The minimum TLS version is 1.2. That is, encryption is valid only between the two sidecars, and between the application container and its own sidecar it is plaintext over loopback inside the Pod.
The default mode is PERMISSIVE, so it accepts both plaintext and mTLS. It is designed for gradually migrating clients that have no sidecar. After migrating everything, you change PeerAuthentication to STRICT to reject plaintext.
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default
namespace: foo
spec:
mtls:
mode: STRICT
To apply it to the whole mesh, put it in the root namespace; per workload, use a selector; and for per-port exceptions, use portLevelMtls. The result in the mTLS migration task document shows this difference — after applying STRICT, only curl coming from the legacy namespace without a sidecar fails. In ambient mode, instead of a sidecar, the node's ztunnel handles mTLS over the HBONE protocol, which is why DISABLE mode isn't supported.
Istio's security best practices state the limits of mTLS clearly. mTLS is authentication, not authorization, so anyone with a valid certificate can reach the service, and to lock it down you must set AuthorizationPolicy in a default-deny pattern. The sidecar intercepts only TCP, so UDP and ICMP pass through, a few ports such as 22 are excluded from inbound capture, and because the application and the sidecar are in the same network and process namespace, the application can delete the redirection rules and bypass the sidecar. So the documentation recommends defense in depth with a NetworkPolicy alongside.
| Comparison | Cilium transparent encryption | Istio mTLS |
|---|---|---|
| Termination point | Node kernel (WireGuard/IPsec) | Sidecar Envoy inside the Pod (node ztunnel in ambient) |
| Application changes | None | None (sidecar injection needed) |
| Identity | Node key (per node) | Workload service account certificate (per workload) |
| Coupling with authorization | Separate from NetworkPolicy | Workload- and request-level authorization possible with AuthorizationPolicy |
| Scope | All L3 traffic between Cilium-managed endpoints | TCP traffic intercepted by the sidecar |
Why upgrading is a security item
According to the patch releases document, patches come out roughly every month, and each minor is supported for about 14 months — after the standard period of 12 months, in the 2-month maintenance mode only vulnerabilities assigned a CVE, dependencies, and core component problems are fixed. The version skew policy says security fixes are backported only to the latest three minor branches. A cluster outside the support window has no fix to receive even when a vulnerability is published, and remaining in that state is itself a defect. The securing a cluster document says to join the kubernetes-announce group to receive security announcements, and the kubeadm certificates document says that moving to the latest patch immediately and staying on a supported minor is how you stay safe.
That the skew policy sets the upgrade order is also tied to security. You can't skip a minor, so a cluster two versions behind must go through two upgrades, and the more you postpone, the greater the cost of catching up. This is why upgrading should be a routine task.
Minimizing the host OS
What the Kubernetes documentation says directly about the node OS concerns boundaries. The securing a cluster document says that because the kubelet allows unauthenticated access by default, production clusters must turn on kubelet authentication and authorization, isolate etcd behind a firewall so only the API server can reach it, restrict Pod access to the cloud metadata API, turn off unused alpha and beta features, and rotate credentials frequently with short lifetimes. The security checklist says not to expose the API server, the kubelet API, or etcd to the internet, and to place workloads of differing sensitivity on separate nodes.
The principle of OS hardening layered on top of that is not to keep anything that isn't needed to run the kubelet, the container runtime, and the CNI. You remove unneeded services and packages — a daemon that isn't running isn't an attack surface even if it has a vulnerability, and a package that isn't installed doesn't need patching. For SSH, according to OpenSSH's sshd_config manual, the default of PermitRootLogin is prohibit-password and the default of PasswordAuthentication is yes, so on nodes the usual setup is to turn off password authentication and leave only key authentication. Kernel modules can be blocked from loading with kernel.modules_disabled from the kernel sysctl documentation, but once you set it to 1, you can neither load nor unload modules and can't undo it, so you turn it on only after loading all the modules you need (those used by the runtime and CNI). The SSH and kernel module items in this section are based not on the Kubernetes documentation but on each tool's official manual, and which packages to remove differs by distribution and CNI, so this section doesn't fix a list.
What it looks like in the field
Turning on STRICT cut off monitoring. If a Prometheus with no sidecar was scraping workloads inside the mesh, the scrape fails along with the switch to STRICT. It is the result of skipping the procedure the Istio migration document describes — finding, with a dashboard, the clients that come in as plaintext while in PERMISSIVE, and locking down per namespace after migrating everything.
Turning on Cilium WireGuard cut off communication between nodes. If UDP 51871 isn't in the cloud security group, the tunnel doesn't come up. For IPsec, it is ESP. Turning on encryption is two lines of Helm values, but if you don't coordinate with the network team on the ports those values require, the entire Pod network stops.
What to check in the next quiz
The quiz asks about the difference in termination point between the two encryption approaches, how same-node traffic is handled, the difference between PERMISSIVE and STRICT, what mTLS can't do on its own, the patch support period, and the nature of modules_disabled.