TT Lab
Get started
Learn Learning paths Courses

Kubernetes Distributions — Build Them Yourself

One line of profile: cis and RKE2 refused to start

Continue in TT Lab

Goal

On a single RKE2 v1.36.4+rke2r1 server in the VM, you read the default configuration, meet with real failure messages what the host must have ready when you change to profile: cis, and check that Pod Security Admission changes from privileged to restricted so that the same privileged Pod is rejected, and then an etcd snapshot as well.

Why it matters

RKE2 is a single-binary distribution in the same family as k3s, but its datastore is the embedded etcd from the start, and it is designed so that security hardening is turned on with a single profile. So what blocks you when you move is not features but policy and host preparation. The cis profile runs etcd as a dedicated user, stops the kubelet from changing kernel values, and enforces the restricted standard on all namespaces. The profile does not satisfy these requirements for you; it only checks them — on an unprepared host it refuses to start at all. Also, if you turn it on later in a cluster that was already running with the default profile, the root etcd left behind when you only stop the service creates a new problem. If you know which defaults change when you move, you can prevent ahead of time the incident where "a Pod that came up until yesterday is rejected today."

Steps

  1. Write to /root/rke2/layout.json version (the version on the first line of rke2 --version, in the form v0.0.0+rke2r0), rke2_binary (the absolute path of the rke2 executable), kubectl (the absolute path of the kubectl that RKE2 unpacked along with it), kubeconfig (the path of the admin kubeconfig), and kubectl_on_default_path (whether kubectl can be found with only the default PATH /usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin, a boolean).
  2. Write to /root/rke2/components.json cni (the name judged from the CNI DaemonSet in kube-system, lowercase), ingress_class (the name of the default IngressClass), datastore (the control plane datastore: etcd or sqlite), storage_classes (a sorted array of the StorageClass names), cluster_cidr, service_cidr, cluster_dns (the CoreDNS Service IP), and overlaps_host (whether either of the two ranges overlaps the host cluster's 10.244.0.0/16 or 10.96.0.0/12, a boolean).
  3. Create a namespace legacy, bring up a privileged Pod priv (image public.ecr.aws/docker/library/busybox:1.37, sleep 86400, privileged: true) with /root/rke2/priv-legacy.yaml, and make it Ready. Then write to /root/rke2/pss-default.json pss_file (the path of the Pod Security Admission configuration file that RKE2 wrote), enforce (the defaults.enforce value of that file), etcd_process_user (the name of the user currently running the etcd process), and priv_uid (the UID of the Pod priv).
  4. Put profile: cis in /etc/rancher/rke2/config.yaml and run systemctl restart rke2-server. If it fails, save one level=fatal line from the rke2-server journal as is to /root/rke2/cis-fatal.txt, stop it with systemctl stop rke2-server so that the retry loop does not run until the host preparation is finished, and clean up with rke2-killall.sh the static Pods that keep running even when you stop the service (including the etcd running as root).
  5. As in the documentation's host requirements, create a system user and group etcd (no home, shell nologin), copy the rke2-cis-sysctl.conf that RKE2 unpacked along with it to /etc/sysctl.d/60-rke2-cis.conf, and apply it (sysctl -p <파일>, where the placeholder is that file, instead of restarting systemd-sysctl). This cluster was created with the default profile, so before starting, check the owner of /var/lib/rancher/rke2/server/db/etcd/member/snap/db and write etcd_uid, etcd_gid, and db_owner_before (사용자:그룹, that is, user:group) to /root/rke2/host-prep.json, and then, without touching the ownership, start the service with systemctl start --no-block rke2-server.
  6. Wait until the API server returns ok on /readyz, and then write to /root/rke2/cis-state.json etcd_process_user, enforce (the current defaults.enforce in rke2-pss.yaml), exempt_namespaces (a sorted array of exemptions.namespaces in that file), db_owner_after (the current 사용자:그룹 of member/snap/db, that is, user:group), netpol_namespaces (a sorted array of the namespaces that have at least one NetworkPolicy), and protect_kernel_defaults (whether protectKernelDefaults: true is in a file of the configuration directory the kubelet reads with --config-dir, a boolean).
  7. Create a namespace shop, apply the same privileged Pod as in step 3 with /root/rke2/priv-shop.yaml (only the namespace is shop), and save the entire rejection output to /root/rke2/denied.txt. Then bring up in shop, with /root/rke2/ok.yaml, a Pod ok that satisfies the restricted standard (same image, sleep 86400) and make it Ready. Do not delete priv in legacy.
  8. Create a snapshot with rke2 etcd-snapshot save --name before-shop and write to /root/rke2/snapshot.json name (the full snapshot name RKE2 gave it), path (the absolute path of the file), size (in bytes, a number), and sha256 (the file hash).
  9. Write to /root/rke2/report.json profile (the value in config.yaml), enforce_before, enforce_after, start_blocker (what the fatal in step 4 demanded: one of etcd-user, sysctl, selinux), legacy_priv_running (whether priv in legacy is Running now, a boolean), snapshot_dir (the directory where the snapshot is stored), and default_sa_automount_disabled (whether automountServiceAccountToken: false is set on the default ServiceAccount of the shop namespace, a boolean).

Notes

What RKE2 put where

Write to /root/rke2/layout.json version (the version on the first line of rke2 --version, in the form v0.0.0+rke2r0), rke2_binary (the absolute path of the rke2 executable), kubectl (the absolute path of the kubectl that RKE2 unpacked along with it), kubeconfig (the path of the admin kubeconfig), and kubectl_on_default_path (whether kubectl can be found with only the default PATH /usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin, a boolean).

k3s installs kubectl in /usr/local/bin along with itself, but RKE2 puts kubectl, crictl, and ctr in bin under the data directory. Check whether it can be found with only the default PATH by using env -i PATH=... bash -c 'command -v kubectl'.

What came in by default

Write to /root/rke2/components.json cni (the name judged from the CNI DaemonSet in kube-system, lowercase), ingress_class (the name of the default IngressClass), datastore (the control plane datastore: etcd or sqlite), storage_classes (a sorted array of the StorageClass names), cluster_cidr, service_cidr, cluster_dns (the CoreDNS Service IP), and overlaps_host (whether either of the two ranges overlaps the host cluster's 10.244.0.0/16 or 10.96.0.0/12, a boolean).

The static Pod manifests are in /var/lib/rancher/rke2/agent/pod-manifests. Read the --service-cluster-ip-range of kube-apiserver and the --cluster-cidr of kube-controller-manager. You can compute the CIDR overlap with Python's ipaddress overlaps.

Under the default profile, a privileged Pod comes up

Create a namespace legacy, bring up a privileged Pod priv (image public.ecr.aws/docker/library/busybox:1.37, sleep 86400, privileged: true) with /root/rke2/priv-legacy.yaml, and make it Ready. Then write to /root/rke2/pss-default.json pss_file (the path of the Pod Security Admission configuration file that RKE2 wrote), enforce (the defaults.enforce value of that file), etcd_process_user (the name of the user currently running the etcd process), and priv_uid (the UID of the Pod priv).

RKE2 makes the PSA configuration into a file under /etc/rancher/rke2 and passes it to the API server. You can see the process user with ps -eo user,comm. Do not delete this Pod even after you change the profile later.

RKE2 refused to start over the single line profile: cis

Put profile: cis in /etc/rancher/rke2/config.yaml and run systemctl restart rke2-server. If it fails, save one level=fatal line from the rke2-server journal as is to /root/rke2/cis-fatal.txt, stop it with systemctl stop rke2-server so that the retry loop does not run until the host preparation is finished, and clean up with rke2-killall.sh the static Pods that keep running even when you stop the service (including the etcd running as root).

The configuration file does not exist by default, so you create it yourself. The service has Restart=always, so it retries every 5 seconds. The fatal message tells you which of the things the cis profile requires of the host is missing. systemctl stop stops only the rke2 process and leaves the containers, so the old etcd can keep running as root and touch files. The cleanup script is in the bin of the install prefix.

Prepare the etcd user and kernel values and start again

As in the documentation's host requirements, create a system user and group etcd (no home, shell nologin), copy the rke2-cis-sysctl.conf that RKE2 unpacked along with it to /etc/sysctl.d/60-rke2-cis.conf, and apply it (sysctl -p <파일>, where the placeholder is that file, instead of restarting systemd-sysctl). This cluster was created with the default profile, so before starting, check the owner of /var/lib/rancher/rke2/server/db/etcd/member/snap/db and write etcd_uid, etcd_gid, and db_owner_before (사용자:그룹, that is, user:group) to /root/rke2/host-prep.json, and then, without touching the ownership, start the service with systemctl start --no-block rke2-server.

The sysctl file of a tarball install is under share/rke2 of the install prefix. The documentation warns that restarting systemd-sysctl on a running cluster can collide with the kernel values the CNI created and cause side effects. The hardening guide explains that the cis profile makes the etcd data directory owned by etcd at startup — what happens to the files the default profile created as root you will check in the next step. If you did not clean up even the static Pods in step 4, a problem arises here.

Check the cluster that came back up with the cis profile

Wait until the API server returns ok on /readyz, and then write to /root/rke2/cis-state.json etcd_process_user, enforce (the current defaults.enforce in rke2-pss.yaml), exempt_namespaces (a sorted array of exemptions.namespaces in that file), db_owner_after (the current 사용자:그룹 of member/snap/db, that is, user:group), netpol_namespaces (a sorted array of the namespaces that have at least one NetworkPolicy), and protect_kernel_defaults (whether protectKernelDefaults: true is in a file of the configuration directory the kubelet reads with --config-dir, a boolean).

Right after the restart the API server refuses connections for a while, so wait by repeating kubectl get --raw /readyz. If it never comes back, look at the etcd container log (/var/log/pods/kube-system_etcd-*). The kubelet in this version receives protect-kernel-defaults not as a command-line flag but through a configuration file (measured) — find --config-dir in ps -eo args and read that directory. With the cis profile, RKE2 itself writes the PSA configuration file and the NetworkPolicies of the default namespaces, but the NetworkPolicies appear a little late after the API is ready, so if the list looks empty, read it again after a moment.

The same manifest was rejected this time

Create a namespace shop, apply the same privileged Pod as in step 3 with /root/rke2/priv-shop.yaml (only the namespace is shop), and save the entire rejection output to /root/rke2/denied.txt. Then bring up in shop, with /root/rke2/ok.yaml, a Pod ok that satisfies the restricted standard (same image, sleep 86400) and make it Ready. Do not delete priv in legacy.

The rejection message lists every field that violated restricted. That list is the list of things to fix (runAsNonRoot, seccompProfile, allowPrivilegeEscalation, capabilities). PSA checks only creation requests, so Pods that were already running are left as they are. Right after you create the namespace, the default ServiceAccount does not exist yet, so it may be rejected for a different reason; check that the reason for the rejection is PodSecurity.

An etcd snapshot after changing the profile

Create a snapshot with rke2 etcd-snapshot save --name before-shop and write to /root/rke2/snapshot.json name (the full snapshot name RKE2 gave it), path (the absolute path of the file), size (in bytes, a number), and sha256 (the file hash).

RKE2 appends the node name and the Unix time after the name you pass. Both rke2 etcd-snapshot ls and kubectl get etcdsnapshotfile show the location. One warning line about the profile key in config.yaml appears, but it has nothing to do with the snapshot.

Summarize as a checklist to verify before moving

Write to /root/rke2/report.json profile (the value in config.yaml), enforce_before, enforce_after, start_blocker (what the fatal in step 4 demanded: one of etcd-user, sysctl, selinux), legacy_priv_running (whether priv in legacy is Running now, a boolean), snapshot_dir (the directory where the snapshot is stored), and default_sa_automount_disabled (whether automountServiceAccountToken: false is set on the default ServiceAccount of the shop namespace, a boolean).

Write based on the record files from earlier steps and the current cluster. The hardening guide explains that RKE2 itself fixes the default ServiceAccount of system namespaces, but namespaces the operator creates are the operator's job — check this for real in shop