Kubernetes Distributions — Build Them Yourself
One line of profile: cis and RKE2 refused to start
Goal
On a single RKE2 v1.36.4+rke2r1 server in the VM, you read the default configuration, meet with real failure messages what the host must have ready when you change to profile: cis,
and check that Pod Security Admission changes from privileged to restricted so that the same privileged Pod is rejected, and then an etcd snapshot as well.
Why it matters
RKE2 is a single-binary distribution in the same family as k3s, but its datastore is the embedded etcd from the start, and it is designed so that security hardening is turned on with a single profile. So what blocks you when you move is not features but policy and host preparation. The cis profile runs etcd as a dedicated user, stops the kubelet from changing kernel values, and enforces the restricted standard on all namespaces. The profile does not satisfy these requirements for you; it only checks them — on an unprepared host it refuses to start at all. Also, if you turn it on later in a cluster that was already running with the default profile, the root etcd left behind when you only stop the service creates a new problem. If you know which defaults change when you move, you can prevent ahead of time the incident where "a Pod that came up until yesterday is rejected today."
Steps
- Write to
/root/rke2/layout.jsonversion(the version on the first line ofrke2 --version, in the form v0.0.0+rke2r0),rke2_binary(the absolute path of the rke2 executable),kubectl(the absolute path of the kubectl that RKE2 unpacked along with it),kubeconfig(the path of the admin kubeconfig), andkubectl_on_default_path(whether kubectl can be found with only the default PATH/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin, a boolean). - Write to
/root/rke2/components.jsoncni(the name judged from the CNI DaemonSet in kube-system, lowercase),ingress_class(the name of the default IngressClass),datastore(the control plane datastore:etcdorsqlite),storage_classes(a sorted array of the StorageClass names),cluster_cidr,service_cidr,cluster_dns(the CoreDNS Service IP), andoverlaps_host(whether either of the two ranges overlaps the host cluster's 10.244.0.0/16 or 10.96.0.0/12, a boolean). - Create a namespace
legacy, bring up a privileged Podpriv(imagepublic.ecr.aws/docker/library/busybox:1.37,sleep 86400,privileged: true) with/root/rke2/priv-legacy.yaml, and make it Ready. Then write to/root/rke2/pss-default.jsonpss_file(the path of the Pod Security Admission configuration file that RKE2 wrote),enforce(the defaults.enforce value of that file),etcd_process_user(the name of the user currently running the etcd process), andpriv_uid(the UID of the Pod priv). - Put
profile: cisin/etc/rancher/rke2/config.yamland runsystemctl restart rke2-server. If it fails, save onelevel=fatalline from the rke2-server journal as is to/root/rke2/cis-fatal.txt, stop it withsystemctl stop rke2-serverso that the retry loop does not run until the host preparation is finished, and clean up withrke2-killall.shthe static Pods that keep running even when you stop the service (including the etcd running as root). - As in the documentation's host requirements, create a system user and group
etcd(no home, shell nologin), copy therke2-cis-sysctl.confthat RKE2 unpacked along with it to/etc/sysctl.d/60-rke2-cis.conf, and apply it (sysctl -p <파일>, where the placeholder is that file, instead of restarting systemd-sysctl). This cluster was created with the default profile, so before starting, check the owner of/var/lib/rancher/rke2/server/db/etcd/member/snap/dband writeetcd_uid,etcd_gid, anddb_owner_before(사용자:그룹, that is, user:group) to/root/rke2/host-prep.json, and then, without touching the ownership, start the service withsystemctl start --no-block rke2-server. - Wait until the API server returns ok on
/readyz, and then write to/root/rke2/cis-state.jsonetcd_process_user,enforce(the current defaults.enforce in rke2-pss.yaml),exempt_namespaces(a sorted array of exemptions.namespaces in that file),db_owner_after(the current사용자:그룹of member/snap/db, that is, user:group),netpol_namespaces(a sorted array of the namespaces that have at least one NetworkPolicy), andprotect_kernel_defaults(whetherprotectKernelDefaults: trueis in a file of the configuration directory the kubelet reads with--config-dir, a boolean). - Create a namespace
shop, apply the same privileged Pod as in step 3 with/root/rke2/priv-shop.yaml(only the namespace is shop), and save the entire rejection output to/root/rke2/denied.txt. Then bring up in shop, with/root/rke2/ok.yaml, a Podokthat satisfies the restricted standard (same image,sleep 86400) and make it Ready. Do not delete priv in legacy. - Create a snapshot with
rke2 etcd-snapshot save --name before-shopand write to/root/rke2/snapshot.jsonname(the full snapshot name RKE2 gave it),path(the absolute path of the file),size(in bytes, a number), andsha256(the file hash). - Write to
/root/rke2/report.jsonprofile(the value in config.yaml),enforce_before,enforce_after,start_blocker(what the fatal in step 4 demanded: one ofetcd-user,sysctl,selinux),legacy_priv_running(whether priv in legacy is Running now, a boolean),snapshot_dir(the directory where the snapshot is stored), anddefault_sa_automount_disabled(whether automountServiceAccountToken: false is set on the default ServiceAccount of the shop namespace, a boolean).
Notes
- In the login shell,
KUBECONFIGand/var/lib/rancher/rke2/binare in the PATH. The grader uses the paths directly without that setting. - Service log:
journalctl -u rke2-server --no-pager | tail - After you clean up in step 4, the API does not respond until it comes back up in step 6. If you only
systemctl stop, the static Pods keep running and the responses look mixed. - Common mistake: running
systemctl restart systemd-sysctlwhile doing the host preparation. The documentation warns of side effects on a running cluster. - Common mistake: only running
systemctl stopand starting again with the static Pods left behind. The old root etcd receives the defragmentation and leaves the db root-owned, and the new etcd that came up as the etcd user cannot open the db and panics (measured). - RKE2 CIS hardening guide · Pod Security Standards · Backup and restore
What RKE2 put where
Write to /root/rke2/layout.json version (the version on the first line of rke2 --version, in the form v0.0.0+rke2r0), rke2_binary (the absolute path of the rke2 executable), kubectl (the absolute path of the kubectl that RKE2 unpacked along with it), kubeconfig (the path of the admin kubeconfig), and kubectl_on_default_path (whether kubectl can be found with only the default PATH /usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin, a boolean).
k3s installs kubectl in /usr/local/bin along with itself, but RKE2 puts kubectl, crictl, and ctr in bin under the data directory. Check whether it can be found with only the default PATH by using env -i PATH=... bash -c 'command -v kubectl'.
What came in by default
Write to /root/rke2/components.json cni (the name judged from the CNI DaemonSet in kube-system, lowercase), ingress_class (the name of the default IngressClass), datastore (the control plane datastore: etcd or sqlite), storage_classes (a sorted array of the StorageClass names), cluster_cidr, service_cidr, cluster_dns (the CoreDNS Service IP), and overlaps_host (whether either of the two ranges overlaps the host cluster's 10.244.0.0/16 or 10.96.0.0/12, a boolean).
The static Pod manifests are in /var/lib/rancher/rke2/agent/pod-manifests. Read the --service-cluster-ip-range of kube-apiserver and the --cluster-cidr of kube-controller-manager. You can compute the CIDR overlap with Python's ipaddress overlaps.
Under the default profile, a privileged Pod comes up
Create a namespace legacy, bring up a privileged Pod priv (image public.ecr.aws/docker/library/busybox:1.37, sleep 86400, privileged: true) with /root/rke2/priv-legacy.yaml, and make it Ready. Then write to /root/rke2/pss-default.json pss_file (the path of the Pod Security Admission configuration file that RKE2 wrote), enforce (the defaults.enforce value of that file), etcd_process_user (the name of the user currently running the etcd process), and priv_uid (the UID of the Pod priv).
RKE2 makes the PSA configuration into a file under /etc/rancher/rke2 and passes it to the API server. You can see the process user with ps -eo user,comm. Do not delete this Pod even after you change the profile later.
RKE2 refused to start over the single line profile: cis
Put profile: cis in /etc/rancher/rke2/config.yaml and run systemctl restart rke2-server. If it fails, save one level=fatal line from the rke2-server journal as is to /root/rke2/cis-fatal.txt, stop it with systemctl stop rke2-server so that the retry loop does not run until the host preparation is finished, and clean up with rke2-killall.sh the static Pods that keep running even when you stop the service (including the etcd running as root).
The configuration file does not exist by default, so you create it yourself. The service has Restart=always, so it retries every 5 seconds. The fatal message tells you which of the things the cis profile requires of the host is missing. systemctl stop stops only the rke2 process and leaves the containers, so the old etcd can keep running as root and touch files. The cleanup script is in the bin of the install prefix.
Prepare the etcd user and kernel values and start again
As in the documentation's host requirements, create a system user and group etcd (no home, shell nologin), copy the rke2-cis-sysctl.conf that RKE2 unpacked along with it to /etc/sysctl.d/60-rke2-cis.conf, and apply it (sysctl -p <파일>, where the placeholder is that file, instead of restarting systemd-sysctl). This cluster was created with the default profile, so before starting, check the owner of /var/lib/rancher/rke2/server/db/etcd/member/snap/db and write etcd_uid, etcd_gid, and db_owner_before (사용자:그룹, that is, user:group) to /root/rke2/host-prep.json, and then, without touching the ownership, start the service with systemctl start --no-block rke2-server.
The sysctl file of a tarball install is under share/rke2 of the install prefix. The documentation warns that restarting systemd-sysctl on a running cluster can collide with the kernel values the CNI created and cause side effects. The hardening guide explains that the cis profile makes the etcd data directory owned by etcd at startup — what happens to the files the default profile created as root you will check in the next step. If you did not clean up even the static Pods in step 4, a problem arises here.
Check the cluster that came back up with the cis profile
Wait until the API server returns ok on /readyz, and then write to /root/rke2/cis-state.json etcd_process_user, enforce (the current defaults.enforce in rke2-pss.yaml), exempt_namespaces (a sorted array of exemptions.namespaces in that file), db_owner_after (the current 사용자:그룹 of member/snap/db, that is, user:group), netpol_namespaces (a sorted array of the namespaces that have at least one NetworkPolicy), and protect_kernel_defaults (whether protectKernelDefaults: true is in a file of the configuration directory the kubelet reads with --config-dir, a boolean).
Right after the restart the API server refuses connections for a while, so wait by repeating kubectl get --raw /readyz. If it never comes back, look at the etcd container log (/var/log/pods/kube-system_etcd-*). The kubelet in this version receives protect-kernel-defaults not as a command-line flag but through a configuration file (measured) — find --config-dir in ps -eo args and read that directory. With the cis profile, RKE2 itself writes the PSA configuration file and the NetworkPolicies of the default namespaces, but the NetworkPolicies appear a little late after the API is ready, so if the list looks empty, read it again after a moment.
The same manifest was rejected this time
Create a namespace shop, apply the same privileged Pod as in step 3 with /root/rke2/priv-shop.yaml (only the namespace is shop), and save the entire rejection output to /root/rke2/denied.txt. Then bring up in shop, with /root/rke2/ok.yaml, a Pod ok that satisfies the restricted standard (same image, sleep 86400) and make it Ready. Do not delete priv in legacy.
The rejection message lists every field that violated restricted. That list is the list of things to fix (runAsNonRoot, seccompProfile, allowPrivilegeEscalation, capabilities). PSA checks only creation requests, so Pods that were already running are left as they are. Right after you create the namespace, the default ServiceAccount does not exist yet, so it may be rejected for a different reason; check that the reason for the rejection is PodSecurity.
An etcd snapshot after changing the profile
Create a snapshot with rke2 etcd-snapshot save --name before-shop and write to /root/rke2/snapshot.json name (the full snapshot name RKE2 gave it), path (the absolute path of the file), size (in bytes, a number), and sha256 (the file hash).
RKE2 appends the node name and the Unix time after the name you pass. Both rke2 etcd-snapshot ls and kubectl get etcdsnapshotfile show the location. One warning line about the profile key in config.yaml appears, but it has nothing to do with the snapshot.
Summarize as a checklist to verify before moving
Write to /root/rke2/report.json profile (the value in config.yaml), enforce_before, enforce_after, start_blocker (what the fatal in step 4 demanded: one of etcd-user, sysctl, selinux), legacy_priv_running (whether priv in legacy is Running now, a boolean), snapshot_dir (the directory where the snapshot is stored), and default_sa_automount_disabled (whether automountServiceAccountToken: false is set on the default ServiceAccount of the shop namespace, a boolean).
Write based on the record files from earlier steps and the current cluster. The hardening guide explains that RKE2 itself fixes the default ServiceAccount of system namespaces, but namespaces the operator creates are the operator's job — check this for real in shop