TT Lab
Get started
Learn Learning paths Courses

Kubernetes Distributions — Build Them Yourself

We moved to OpenShift and it died with permission denied

Continue in TT Lab

This lab imitates OpenShift on k3s, not on OpenShift itself

OpenShift requires at least 8 vCPU, 16GB of memory, and 120GB of storage even when installed on a single node (OCP 4.21 documentation), so it cannot be run on this VM (8GiB). So on the k3s v1.36.4+k3s1 in the VM, you create the same symptoms by putting in by hand the values that OpenShift's restricted-v2 SCC would fill in for the Pod (an arbitrary UID from the range, the root group, fsGroup). k3s knows neither SCCs nor project annotations, so the annotations are used only as the basis for calculation — in the last step you also confirm that fact.

Goal

You reproduce an image that writes to a root-owned directory dying with permission denied under an arbitrary UID, fix the image with the root group and g=u so that it comes up with any UID in the range, and distinguish what each of fsGroup, a root initContainer, and a read-only root does and cannot do for this problem.

Why it matters

Most of the reasons why an image that ran fine on Kubernetes first dies on OpenShift come down to the user. OpenShift does not trust the image's USER and runs the container with a large UID assigned per project. The design is meant so that even if there is a container escape vulnerability, the process does not become a meaningful user on the host. You cannot know that user in advance, so you cannot match it to the "owner" of the files in the image, and instead the official guidance is to give permissions to the root group that it always belongs to. Attempts to work around it with Pod settings (fsGroup, an initContainer that chowns as root) are mostly blocked by policy or have no effect on image directories.

Steps

  1. Create a namespace ocp-sim and attach the labels pod-security.kubernetes.io/enforce=restricted and pod-security.kubernetes.io/warn=restricted and the annotations openshift.io/sa.scc.uid-range=1000680000/10000 and openshift.io/sa.scc.supplemental-groups=1000680000/10000. Then write to /root/ocp/project.json uid_range (the annotation value as is), default_uid (the UID that restricted-v2's MustRunAsRange would pick by default), and max_uid (the last UID of the range).
  2. Create a Pod legacy with /root/ocp/legacy.yaml. Image localhost/ocp-app:legacy (preloaded in the VM); the Pod securityContext is runAsNonRoot true, runAsUser the default_uid from step 1, runAsGroup 0, fsGroup 1000680000, and seccompProfile RuntimeDefault; the container is allowPrivilegeEscalation false, capabilities drop ALL, and terminationMessagePolicy: FallbackToLogsOnError. Once a restart has happened at least once, write to /root/ocp/crash.json uid, restart_count, and error (the line containing Permission denied in the termination message, as is).
  3. With the legacy image, create a Pod probe that only does sleep 86400, with the same securityContext as in step 2, in /root/ocp/probe.yaml, and leave it Ready. Investigate the identity inside it and write to /root/ocp/identity.json uid, gid, groups (a sorted array of the numbers from id -G), whoami_ok (whether whoami succeeds), home (the value of the HOME environment variable), and home_writable (whether you can write to that HOME).
  4. Write /root/ocp/app/Containerfile so that, with the same base and script as the legacy image, it makes /app owned by the root group (GID 0), sets the group permissions equal to the owner permissions (g=u), and then specifies a numeric USER (a value other than 0). Build localhost/ocp-app:fixed with buildah and load it into k3s's containerd (k8s.io namespace), and then write to /root/ocp/build.json tool, image, image_id (the status.id of k3s crictl inspecti), and user (the User in the image configuration).
  5. With the fixed image, create Pods fixed-a (runAsUser is default_uid) and fixed-b (runAsUser is max_uid) with the same securityContext as in step 2, in /root/ocp/fixed-a.yaml and /root/ocp/fixed-b.yaml, and leave both Ready. Write to /root/ocp/arbitrary.json, keyed by Pod name, uid (id -u inside the container) and data (the 소유자UID:그룹GID:권한8진수 of /app/data, in the form 0:0:755, where the placeholders mean owner UID, group GID, and permission in octal).
  6. Test two ways of holding on without fixing the legacy image. (1) With /root/ocp/legacy-emptydir.yaml, create a Pod legacy-ed (same as step 2 but with the emptyDir data mounted on /app/data) and leave it Ready. (2) With /root/ocp/legacy-init.yaml, apply a Pod legacy-init in which a root (runAsUser 0) initContainer fix-perms chowns /app/data, and save the output to /root/ocp/init-denied.txt. Write to /root/ocp/alternatives.json fsgroup_fixed_image_dir (whether the legacy Pod that has an fsGroup succeeded in writing to the image directory), emptydir_data (the 그룹GID:권한8진수 of /app/data inside legacy-ed, where the placeholders mean group GID and permission in octal), and root_init_admitted (whether legacy-init was admitted).
  7. Create a Pod ro-bare that adds only readOnlyRootFilesystem: true to the fixed image, with /root/ocp/ro-bare.yaml, and see it fail; then, with /root/ocp/fixed-ro.yaml, create a Pod fixed-ro with the same settings plus emptyDirs mounted on /app/data (named data) and /tmp (named tmp) and the environment variable HOME=/tmp, and leave it Ready. Write to /root/ocp/readonly.json bare_error (the line containing Read-only in the ro-bare termination message, as is), home (HOME inside fixed-ro), home_writable, and etc_writable (whether you can write to /etc inside fixed-ro).
  8. Write to /root/ocp/report.json root_cause (the cause of the failure in step 2: one of image-dir-owner, missing-capability, selinux), whoami_ok_here (whether whoami works with an arbitrary UID on this k3s), openshift_runtime_adds_passwd_entry (whether the OpenShift documentation says CRI-O puts the arbitrary UID into /etc/passwd), fsgroup_fixes_image_dir, k3s_applies_uid_range (whether this k3s fills in the annotation's UID for a Pod with no runAsUser), uid_range_default (the default UID calculated from the annotation), and running_uids (a sorted array of the UIDs of fixed-a and fixed-b, which are Ready now).

Notes

Dress up a namespace like an OpenShift project

Create a namespace ocp-sim and attach the labels pod-security.kubernetes.io/enforce=restricted and pod-security.kubernetes.io/warn=restricted and the annotations openshift.io/sa.scc.uid-range=1000680000/10000 and openshift.io/sa.scc.supplemental-groups=1000680000/10000. Then write to /root/ocp/project.json uid_range (the annotation value as is), default_uid (the UID that restricted-v2's MustRunAsRange would pick by default), and max_uid (the last UID of the range).

The uid-range annotation accepts only a single <시작>/<길이> block (start/length). MustRunAsRange uses the minimum of the range as the default. The k3s in this VM does not read this annotation, so in later steps you put that value into the Pod directly.

You moved it to OpenShift and it died with permission denied

Create a Pod legacy with /root/ocp/legacy.yaml. Image localhost/ocp-app:legacy (preloaded in the VM); the Pod securityContext is runAsNonRoot true, runAsUser the default_uid from step 1, runAsGroup 0, fsGroup 1000680000, and seccompProfile RuntimeDefault; the container is allowPrivilegeEscalation false, capabilities drop ALL, and terminationMessagePolicy: FallbackToLogsOnError. Once a restart has happened at least once, write to /root/ocp/crash.json uid, restart_count, and error (the line containing Permission denied in the termination message, as is).

The values the restricted-v2 SCC would fill in on OpenShift are put in by hand here. The Containerfile of the legacy image is in /root/ocp/app/Containerfile.legacy. If you use FallbackToLogsOnError, the tail of the log is left as the termination message in the Pod status, so it does not disappear even after the container comes up again.

What it means to live as a user that is not in /etc/passwd

With the legacy image, create a Pod probe that only does sleep 86400, with the same securityContext as in step 2, in /root/ocp/probe.yaml, and leave it Ready. Investigate the identity inside it and write to /root/ocp/identity.json uid, gid, groups (a sorted array of the numbers from id -G), whoami_ok (whether whoami succeeds), home (the value of the HOME environment variable), and home_writable (whether you can write to that HOME).

containerd does not block starting a container with a UID that is not in the image's /etc/passwd, but it cannot tell the name and home directory. According to the documentation, OpenShift's CRI-O puts an entry for the arbitrary UID into /etc/passwd — what you see here is that difference. Check writability with test -w.

Fix it on the image side: root group and g=u

Write /root/ocp/app/Containerfile so that, with the same base and script as the legacy image, it makes /app owned by the root group (GID 0), sets the group permissions equal to the owner permissions (g=u), and then specifies a numeric USER (a value other than 0). Build localhost/ocp-app:fixed with buildah and load it into k3s's containerd (k8s.io namespace), and then write to /root/ocp/build.json tool, image, image_id (the status.id of k3s crictl inspecti), and user (the User in the image configuration).

An image built by buildah exists only in the buildah store. Take it out with buildah push <이미지> docker-archive:<파일>:<이름> and then load it with k3s ctr -n k8s.io images import <파일> (the placeholders are the image, the file, and the name). The executable also needs group execute permission for an arbitrary UID to run it.

Does it come up with any UID in the range

With the fixed image, create Pods fixed-a (runAsUser is default_uid) and fixed-b (runAsUser is max_uid) with the same securityContext as in step 2, in /root/ocp/fixed-a.yaml and /root/ocp/fixed-b.yaml, and leave both Ready. Write to /root/ocp/arbitrary.json, keyed by Pod name, uid (id -u inside the container) and data (the 소유자UID:그룹GID:권한8진수 of /app/data, in the form 0:0:755, where the placeholders mean owner UID, group GID, and permission in octal).

OpenShift runs the same image with a different UID for each project. An image that works with only one specific UID breaks again the moment you move it. Look with stat -c '%u:%g:%a'.

Why fsGroup and a root initContainer are not the answer

Test two ways of holding on without fixing the legacy image. (1) With /root/ocp/legacy-emptydir.yaml, create a Pod legacy-ed (same as step 2 but with the emptyDir data mounted on /app/data) and leave it Ready. (2) With /root/ocp/legacy-init.yaml, apply a Pod legacy-init in which a root (runAsUser 0) initContainer fix-perms chowns /app/data, and save the output to /root/ocp/init-denied.txt. Write to /root/ocp/alternatives.json fsgroup_fixed_image_dir (whether the legacy Pod that has an fsGroup succeeded in writing to the image directory), emptydir_data (the 그룹GID:권한8진수 of /app/data inside legacy-ed, where the placeholders mean group GID and permission in octal), and root_init_admitted (whether legacy-init was admitted).

fsGroup is a mechanism that changes the group ownership of volumes attached to the Pod, and it does not change the image filesystem. An emptyDir can be written to, but it disappears along with the Pod. Neither the restricted standard nor the restricted-v2 SCC allows UID 0.

Open up only the places you write on a read-only root

Create a Pod ro-bare that adds only readOnlyRootFilesystem: true to the fixed image, with /root/ocp/ro-bare.yaml, and see it fail; then, with /root/ocp/fixed-ro.yaml, create a Pod fixed-ro with the same settings plus emptyDirs mounted on /app/data (named data) and /tmp (named tmp) and the environment variable HOME=/tmp, and leave it Ready. Write to /root/ocp/readonly.json bare_error (the line containing Read-only in the ro-bare termination message, as is), home (HOME inside fixed-ro), home_writable, and etc_writable (whether you can write to /etc inside fixed-ro).

A read-only root reveals where the image writes. Give a volume to every path it writes, and point places that tools use by default, such as HOME, to a writable path. Do not delete ro-bare; leave it as evidence.

A checklist for inspecting the image before moving

Write to /root/ocp/report.json root_cause (the cause of the failure in step 2: one of image-dir-owner, missing-capability, selinux), whoami_ok_here (whether whoami works with an arbitrary UID on this k3s), openshift_runtime_adds_passwd_entry (whether the OpenShift documentation says CRI-O puts the arbitrary UID into /etc/passwd), fsgroup_fixes_image_dir, k3s_applies_uid_range (whether this k3s fills in the annotation's UID for a Pod with no runAsUser), uid_range_default (the default UID calculated from the annotation), and running_uids (a sorted array of the UIDs of fixed-a and fixed-b, which are Ready now).

You can check whether k3s uses the annotation by sending a Pod without runAsUser with --dry-run=server -o json and seeing whether runAsUser was filled in the returned spec. The grader also recalculates using the same method and the records from earlier steps.