Kubernetes Distributions — Build Them Yourself
RKE2 — a distribution whose defaults lean toward security
One-line summary
RKE2 is set up from a single binary like k3s, but its datastore is etcd from the start, and it is a distribution in which the single line profile: cis turns on hardening such as restricted Pod security and a dedicated etcd user. That one line turns on only when the host is prepared.
Why this was needed
k3s is a distribution made lightweight for edge and development environments, so its datastore default is sqlite and its components come in on the convenient side. But in front of a customer who brings a CIS benchmark checklist, such as a public institution or a financial company, the "convenient defaults" themselves become audit items. If a person repeats the work at every inspection — fixing API server flags one by one, matching etcd directory permissions, attaching Pod security labels to each namespace — one spot is bound to be missed.
RKE2 moves that repetition into the distribution. As the RKE2 CIS hardening guide explains, when you turn on profile, RKE2 itself writes the PSA configuration file and the NetworkPolicies of the default namespaces, runs etcd as the etcd user, and narrows file permissions. In return, the host-side preparation (kernel values, the etcd user) is left to the operator, and at startup it checks whether that preparation has been done. The reason the design is "stop if it isn't done" rather than "fix it for you" is that if a distribution quietly changed host settings, it could break other software on that host.
How it works
First, the installation layout differs from k3s. As in the quick start documentation, the kubeconfig is at /etc/rancher/rke2/rke2.yaml, and kubectl, crictl, and ctr are in /var/lib/rancher/rke2/bin, which is not in the default PATH. The control plane is not bundled into one process as in k3s; it is static Pods (etcd, kube-apiserver, and so on) started by the kubelet.
The result of setting it up with the default profile on this VM (measured, v1.36.4+rke2r1):
cni canal (문서의 기본값)
ingress class traefik (v1.36 부터 새 클러스터의 기본, 문서)
datastore etcd 정적 파드
storage class 없음 — k3s 의 local-path 같은 것이 들어오지 않는다
대역 파드 10.42.0.0/16 · 서비스 10.43.0.0/16 · DNS 10.43.0.10
PSA 설정 /etc/rancher/rke2/rke2-pss.yaml, enforce: privileged
etcd 프로세스 root
The CIDR ranges are the same as the defaults in the server configuration reference, and since they do not overlap with the host cluster (10.244/10.96) that runs this lab VM, they are used as they are.
What changes when you turn on profile: cis is summarized in the Pod Security Standards documentation. The same file was rewritten like this (measured).
defaults:
enforce: "restricted"
audit: "restricted"
warn: "restricted"
exemptions:
namespaces: [kube-system, compliance-operator-system, tigera-operator]
In addition, protect-kernel-defaults is turned on for the kubelet, so if kernel values differ from expectations the kubelet stops (in this version it was not a flag but entered as protectKernelDefaults: true in /var/lib/rancher/rke2/agent/etc/kubelet.conf.d/00-rke2-defaults.conf, measured), and NetworkPolicies are created in default, kube-public, and kube-system. RKE2 also turns off the automatic token mounting of the default ServiceAccount in system namespaces (on this VM it was false in default, kube-public, kube-system, and kube-node-lease, measured). But the guide leaves the NetworkPolicies and the default ServiceAccount of namespaces the operator creates later as "operator intervention required." The two namespaces created in the lab had neither NetworkPolicies nor an automount value (measured).
What it looks like in the field
Suppose that, ahead of an inspection, you put profile: cis into a cluster that had been running for months with the default profile and restarted. This is what happened, in order, on this VM (measured).
1) level=fatal msg="missing required: user: unknown user etcd ..." ← 시작 거부
2) systemctl stop → etcd 사용자·sysctl 준비 → 다시 시작
→ etcd 컨테이너: failed to open database .../member/snap/db ... panic
3) 데이터 디렉터리는 etcd:etcd 로 바뀌었는데 member/snap/db 만 root:root
4) rke2-killall.sh 로 남은 컨테이너까지 내리고 다시 시작 → 9초 만에 readyz ok, db 도 etcd:etcd
The cause was that systemctl stop stops only the rke2 process and leaves the static Pod containers. Even after the stop, the etcd running as root was still there (measured), and that old etcd received the Defragmenting etcd database sent by the newly started RKE2 and rewrote the db file as root-owned. Then the new etcd that came up as the etcd user could not open that file. This is the causality as seen from the log timestamps and the file owner. Conversely, after the old containers were cleaned up, RKE2 matched the ownership itself at startup without any chown (measured) — this is what the guide means by "make the etcd data directory owned by etcd."
It was not over even after the API came back. RKE2 wrote the NetworkPolicies 62 seconds after startup, and the default ServiceAccount appeared in the first namespace created after the restart 16 seconds after the namespace was created (43 seconds on another VM) (measured). If you put in a Pod in between, it is rejected not by PodSecurity but with serviceaccount "default" not found, which makes it look as if it were blocked by policy.
Step 2 is especially confusing. The configuration file was fixed correctly everywhere and the service is activating, yet kubectl only gives connection timeouts. The cause is not in the service log but only in the etcd container log (/var/log/pods/kube-system_etcd-*).
On the Pod side, the same manifest as the privileged Pod that came up fine under the default profile is rejected in a new namespace with violates PodSecurity "restricted:latest". But a privileged Pod that was already running stays Running. PSA checks only creation requests, so it is quiet on the day you turn the profile on and suddenly blocked on the day that Pod is redeployed. Workloads moved from k3s are likewise rejected not because it is RKE2 but because it is RKE2 with the cis profile turned on — RKE2 with the default profile accepts them.
One more thing: the installation path also depends on the environment. If /usr/local is read-only or a separate mount point, the install script unpacks into /opt/rke2 (the quick start documentation and the script comments). On this lab VM, /usr/local is a large scratch disk, and the root that has /opt was 2.4GiB and already 89% full (measured, the installed files are 128M). If left as is, the root would nearly fill up, so INSTALL_RKE2_TAR_PREFIX was set explicitly.
What really matters in practice
- Turn on the cis profile when you first set up. The hardening guide also assumes a state where RKE2 has been installed but not yet run. Do the host preparation (the etcd user, the sysctl file) first and give the profile from the first start. If you have to turn it on later, do not just stop the service; bring down even the static Pods with
rke2-killall.shand then start. - Do not restart systemd-sysctl on a running node. The hardening guide warns of side effects that conflict with kernel values set by the CNI. Apply only the one file with
sysctl -p. - PSA does not inspect the past. After turning the profile on, you need a list for putting existing workloads back in with a server dry-run.
- Snapshots are kept by default every 12 hours, 5 of them. Before a big operation such as a profile change, take a separate named snapshot with
rke2 etcd-snapshot save, and, as in the backup and restore documentation, keep the server token (/var/lib/rancher/rke2/server/token) along with it. According to the documentation, the bootstrap data in the snapshot (confidential data such as the CA certificate) is decrypted with that token, so when restoring to a different host, it cannot be used without the token. You also need to check the location. The documentation's flag table gives the default location as${data-dir}/db/snapshots, but on this VM the file was in/var/lib/rancher/rke2/server/db/snapshots(measured).
What you will do in the next lab
On the RKE2 in the VM, you record the install location and the default components, bring up a privileged Pod under the default profile, and then change to profile: cis. You get the startup-refusal message, clean up even the static Pods, prepare the etcd user and sysctl and bring it up again, and check how the ownership changed. Then you confirm that the same manifest is rejected and that the existing Pod remains, take a snapshot, and summarize in a report.