TT Lab
Get started
Learn Learning paths Courses

Kubernetes Distributions — Build Them Yourself

OpenShift's arbitrary UIDs and SCCs — the image has to be ready first

Continue in TT Lab

One-line summary

OpenShift runs containers not as the image's USER but as an arbitrary UID assigned per project, and puts that user in the root group (GID 0), so only images whose writable directories are owned by the root group and group-writable move over as they are.

Why this was needed

In ordinary Kubernetes, an image with no USER runs as root, and one with a USER runs as that user. So many images stand on the assumption that "root writes into a directory root created." OpenShift does not accept this assumption. The OCP 4.21 image creation guidelines state that containers run by default with an arbitrarily assigned UID, and explain the reason: so that even if a process escapes through a container engine vulnerability, it does not gain high privileges on the host.

The UID can differ every time, so the image author cannot know in advance "which user it will run as." So the guidelines base things on the group instead of the owner. The container user is always a member of the root group, so if you leave the directories and files you write to owned by the root group and give the group read and write permission (and execute permission for executables), it works whatever UID comes. The guidelines' example is chgrp -R 0 <디렉터리> && chmod -R g=u <디렉터리> (where the placeholder is the directory). Also, this user has no privileges, so it cannot open ports below 1024.

One thing to state up front. The documentation says OpenShift requires at least 8 vCPU, 16GB RAM, and 120GB of storage even for a single-node installation, so it cannot be run on our 8GiB lab VM. The OpenShift behavior in this text is what was confirmed in the documentation, and the next lab reproduces the symptoms by building the same conditions by hand on k3s.

How it works

It is the SCC (Security Context Constraints) that picks the UID. According to the OCP 4.19 SCC documentation, the SCC that authenticated users use by default is restricted-v2, and it has these properties.

runAsUser        MustRunAsRange   범위를 네임스페이스 주석에서 가져온다
seLinuxContext   MustRunAs        MCS 레벨도 네임스페이스 주석에서
fsGroup          MustRunAs        supplemental-groups 주석, 없으면 uid-range 로
capabilities     ALL 을 떨어뜨림   NET_BIND_SERVICE 만 명시적으로 더할 수 있다
seccomp          runtime/default
allowPrivilegeEscalation  설정하지 않거나 false 여야 한다

The range is the openshift.io/sa.scc.uid-range annotation of the project (namespace), and it accepts a single <시작>/<길이> block (start/length). If the Pod does not request runAsUser, the minimum of the range becomes the default. There are also differences by version. In the 4.19 documentation restricted-v2 is the default, but in the 4.20 and 4.21 SCC documentation, restricted-v3, which enforces user namespaces (hostUsers: false), was added as the default for new installations. Either way, it is the same in that it runs as a UID from the range.

Separately from the SCC, Pod Security Admission also runs. The 4.21 Pod Security Admission documentation explains the two as independent mechanisms. Globally it enforces privileged and uses restricted only for warnings and audits, and the warn and audit labels of a namespace are automatically synchronized to match the SCCs the service account can use. A workload must pass both.

What it looks like in the field

This is the result of putting in by hand, on the k3s VM, the values restricted-v2 would fill (uid 1000680000, gid 0, fsGroup) and running a common legacy image (/app/data created by root, permission 755) (measured).

id: uid=1000680000 gid=0(root) groups=0(root),1000680000
whoami: whoami: unknown uid 1000680000
HOME=/
startup failed: cannot write /app/data/started
/app/run.sh: line 6: can't create /app/data/started: Permission denied

Here I tested, on the same VM, three workarounds that people commonly try (measured).

When I fixed the image with chgrp -R 0 /app && chmod -R g=u /app, it came up with both the first UID and the last UID of the range (1000689999). One more snag came up then. The first build, in which the script file had no execute bit, failed with RunContainerError and exec: "/app/run.sh": permission denied — this is why the guidelines require group execute permission on executables.

The whoami failure and HOME=/ are what was seen on containerd. The OpenShift documentation explains that CRI-O can put the arbitrary UID into the container's /etc/passwd, so on OpenShift the name lookup may work. However, the 4.21 documentation also writes at the same place that if the image contains /etc/passwd, CRI-O may fail that injection and be unable to resolve the running UID, and it warns that opening up the permissions of that file is dangerous. So you must not carry over "whoami does not work in our k3s" as an OpenShift symptom. On the other hand, for tools that write caches to HOME, it is safer to point HOME to a writable path everywhere.

What really matters in practice

What you will do in the next lab

On the k3s VM, you create a namespace that imitates an OpenShift project and record the legacy image dying with permission denied under an arbitrary UID. You investigate the identity, fix the image with buildah and move it to k3s, bring it up with two UIDs, compare fsGroup, emptyDir, and a root initContainer, try a read-only root, and then summarize in a report.