Kubernetes Distributions — Build Them Yourself
OpenShift's arbitrary UIDs and SCCs — the image has to be ready first
One-line summary
OpenShift runs containers not as the image's USER but as an arbitrary UID assigned per project, and puts that user in the root group (GID 0), so only images whose writable directories are owned by the root group and group-writable move over as they are.
Why this was needed
In ordinary Kubernetes, an image with no USER runs as root, and one with a USER runs as that user. So many images stand on the assumption that "root writes into a directory root created." OpenShift does not accept this assumption. The OCP 4.21 image creation guidelines state that containers run by default with an arbitrarily assigned UID, and explain the reason: so that even if a process escapes through a container engine vulnerability, it does not gain high privileges on the host.
The UID can differ every time, so the image author cannot know in advance "which user it will run as." So the guidelines base things on the group instead of the owner. The container user is always a member of the root group, so if you leave the directories and files you write to owned by the root group and give the group read and write permission (and execute permission for executables), it works whatever UID comes. The guidelines' example is chgrp -R 0 <디렉터리> && chmod -R g=u <디렉터리> (where the placeholder is the directory). Also, this user has no privileges, so it cannot open ports below 1024.
One thing to state up front. The documentation says OpenShift requires at least 8 vCPU, 16GB RAM, and 120GB of storage even for a single-node installation, so it cannot be run on our 8GiB lab VM. The OpenShift behavior in this text is what was confirmed in the documentation, and the next lab reproduces the symptoms by building the same conditions by hand on k3s.
How it works
It is the SCC (Security Context Constraints) that picks the UID. According to the OCP 4.19 SCC documentation, the SCC that authenticated users use by default is restricted-v2, and it has these properties.
runAsUser MustRunAsRange 범위를 네임스페이스 주석에서 가져온다
seLinuxContext MustRunAs MCS 레벨도 네임스페이스 주석에서
fsGroup MustRunAs supplemental-groups 주석, 없으면 uid-range 로
capabilities ALL 을 떨어뜨림 NET_BIND_SERVICE 만 명시적으로 더할 수 있다
seccomp runtime/default
allowPrivilegeEscalation 설정하지 않거나 false 여야 한다
The range is the openshift.io/sa.scc.uid-range annotation of the project (namespace), and it accepts a single <시작>/<길이> block (start/length). If the Pod does not request runAsUser, the minimum of the range becomes the default. There are also differences by version. In the 4.19 documentation restricted-v2 is the default, but in the 4.20 and 4.21 SCC documentation, restricted-v3, which enforces user namespaces (hostUsers: false), was added as the default for new installations. Either way, it is the same in that it runs as a UID from the range.
Separately from the SCC, Pod Security Admission also runs. The 4.21 Pod Security Admission documentation explains the two as independent mechanisms. Globally it enforces privileged and uses restricted only for warnings and audits, and the warn and audit labels of a namespace are automatically synchronized to match the SCCs the service account can use. A workload must pass both.
What it looks like in the field
This is the result of putting in by hand, on the k3s VM, the values restricted-v2 would fill (uid 1000680000, gid 0, fsGroup) and running a common legacy image (/app/data created by root, permission 755) (measured).
id: uid=1000680000 gid=0(root) groups=0(root),1000680000
whoami: whoami: unknown uid 1000680000
HOME=/
startup failed: cannot write /app/data/started
/app/run.sh: line 6: can't create /app/data/started: Permission denied
Here I tested, on the same VM, three workarounds that people commonly try (measured).
- Giving only fsGroup: The Pod above already had an fsGroup, but
/app/datainside the image was still root 755. fsGroup is a mechanism that changes the group of volumes attached to the Pod. - Mounting an emptyDir over it: It was created with group 1000680000 and permission 2777, so writing worked. In exchange, when the Pod disappears the data disappears too, and the files the image put at that path are hidden.
- chown with a root initContainer: It was not even created, with
violates PodSecurity "restricted:latest": runAsUser=0.
When I fixed the image with chgrp -R 0 /app && chmod -R g=u /app, it came up with both the first UID and the last UID of the range (1000689999). One more snag came up then. The first build, in which the script file had no execute bit, failed with RunContainerError and exec: "/app/run.sh": permission denied — this is why the guidelines require group execute permission on executables.
The whoami failure and HOME=/ are what was seen on containerd. The OpenShift documentation explains that CRI-O can put the arbitrary UID into the container's /etc/passwd, so on OpenShift the name lookup may work. However, the 4.21 documentation also writes at the same place that if the image contains /etc/passwd, CRI-O may fail that injection and be unable to resolve the running UID, and it warns that opening up the permissions of that file is dangerous. So you must not carry over "whoami does not work in our k3s" as an OpenShift symptom. On the other hand, for tools that write caches to HOME, it is safer to point HOME to a writable path everywhere.
What really matters in practice
- Run the image with an arbitrary UID before moving. Even on an ordinary cluster, if you give runAsUser a large value and runAsGroup 0, most of the permission problems you would hit on OpenShift show up ahead of time. Running it with two different UIDs from the range is the surest way.
- What you fix is the image. Instead of chowning to a specific UID, use the root group and g=u, and write USER as a number (the documentation explains that the build fails if an S2I image has no numeric USER).
- Granting a broad SCC such as anyuid is a last resort. The documentation warns against modifying the default SCCs, and a broad SCC revives the risk that arbitrary UIDs were meant to block.
- If you turn on a read-only root as well, it reveals where the image writes. The list of giving a volume to every path it writes becomes the operations document.
What you will do in the next lab
On the k3s VM, you create a namespace that imitates an OpenShift project and record the legacy image dying with permission denied under an arbitrary UID. You investigate the identity, fix the image with buildah and move it to k3s, bring it up with two UIDs, compare fsGroup, emptyDir, and a root initContainer, try a read-only root, and then summarize in a report.