The Runtime Gate — An Unvalidated String Against What Actually Loaded
Goal
You deal with both sides of the runtime gate together. On the cluster side, you see what a RuntimeClass chooses and what it does not validate, and on the node side, you create two cases in which configuration is silently ignored. Finally, you create yourself a checker that cross-checks the two.
Why it matters
If a Pod attached to a node but dies at container creation, that is the runtime side. The handler of a RuntimeClass is just a string, and the apiserver does not validate it. That is because the containerd configuration can differ per node, so there is originally no entity that does the validating. So a typo gets applied, gets scheduled, and shows up only at the moment that node creates the container.
On the node side there are two more silent failures. If the imports glob is off from the file name, containerd passes without a word, and if the configuration version rises to 3, the plugin name itself changes, so configuration written under the old name is ignored entirely. In both cases the file has the desired contents written in it intact. So a check must look not at the file but at the loaded configuration, and that judgment becomes useful only when it is matched against the cluster's RuntimeClass list. The final deliverable of this lab is that checker.
Environment
This Pod has no real containerd. So the configuration is handled with a TOML parser, and the checker is made to accept a dump as an argument so that the grader actually runs it with two kinds of dump. On the cluster side, the real apiserver and scheduler brought up by kwok mean that RuntimeClass registration and overhead computation are real. The working directory is /root/gpurt, and under it you use k8s/, etc/, bin/, fixtures/, and out/.
Steps
- Create the
gpu-rtnamespace and two RuntimeClasses. - See that a Pod is scheduled even with a mistyped handler and write it in
out/unvalidated.txt. - Create two nodes with cpu 1 and see the result split by the RuntimeClass overhead.
- Create an
importsglob that fails to catch the drop-in and write it inout/imports.txt. - Rewrite the same configuration in version 3 format in
etc/config-v3.toml. - With
bin/rc-crosscheck.sh, create a checker that cross-checks the cluster and the dump. - Create a drop-in that changes the default runtime and write its blast radius in
out/blast.txt. - Imitate a partially loaded dump, run the checker on it, and save the result in
out/crosscheck.txt.
Notes
- TOML builds its hierarchy by table names, not by indentation. Even if it looks fine to the eye, it often attaches to the wrong place, so read it once with a parser to check.
- If the step 6 checker reads
/etc/containerd/config.toml, the grader rejects it. That is because this accident occurs at the spot where the file and what is loaded diverge. - There is a reason to keep the two nodes separate in step 3. If you put them on one node, the Pod that goes in first takes the slot and hides the effect of the overhead.
Create the object that chooses the handler
Create the gpu-rt namespace, and write two RuntimeClasses in /root/gpurt/k8s/runtimeclasses.yaml and apply them. The names and handlers are nvidia and nvidia-cdi respectively.
A RuntimeClass is a cluster-scoped object, and the handler value must match the runtimes.<이름> (the placeholder is the name) in the node's containerd configuration. When kubelet creates a Pod, it carries this string in the CRI request, and containerd looks for the same name in its own configuration. apiVersion is node.k8s.io/v1.
Even a mistyped handler passes the apiserver
In /root/gpurt/k8s/runtimeclass-typo.yaml, write together a RuntimeClass nvidia-typo with a deliberately wrong handler and a Pod probe-typo that uses it, and apply them to gpu-rt. Write the facts you confirm in /root/gpurt/out/unvalidated.txt as three lines. They are API_VALIDATES_HANDLER, FAILS_AT, and VISIBLE_TO.
The apiserver does not validate the handler string. That is because the containerd configuration can differ per node, so there is originally no entity that does the validating. So a typo gets applied and gets scheduled. The spot where the mismatch shows up is the moment that node creates the container, and only the node knows it. For the values of the three lines, write no, then in a single word at which stage it fails, and who knows it.
The RuntimeClass overhead changes scheduling
Create two kwok nodes with cpu 1 (gpu-tight-a and gpu-tight-b) with /root/gpurt/k8s/node-tight.yaml. Then in /root/gpurt/k8s/overhead.yaml write a RuntimeClass nvidia-overhead whose overhead.podFixed is cpu 250m and memory 128Mi, and two Pods that each request 900m of cpu, and apply them. fits is placed on node a without the RuntimeClass, and over is placed on node b with that RuntimeClass attached.
The overhead of a RuntimeClass is the extra resources that runtime consumes for each Pod. The apiserver fills that value into the Pod's spec.overhead, and the scheduler adds it to the container requests to compute the slots. So even for the same 900m request, with overhead attached it becomes 1150m and does not fit on a node with cpu 1. The reason to keep the two nodes separate is that if you put them on one node, the Pod that goes in first takes the slot and blurs the comparison. kwok nodes are managed only if they have the kwok.x-k8s.io/node: fake annotation.
If the imports glob is off, it passes silently
In /root/gpurt/etc/config.toml, create the accident node's main configuration. Put version = 2, under the CRI plugin default_runtime_name = "runc" and runtimes.runc, and an imports glob that fails to catch the drop-in. Create the drop-in as runtimes.nvidia in /root/gpurt/etc/conf.d/99-nvidia.toml. Write the confirmation result in /root/gpurt/out/imports.txt as four lines.
A glob matching nothing is not an error but a normal result, so containerd does not even leave a log. So if you write the extension as .conf and the file ends up as .toml, that configuration is never read. The four lines of imports.txt are GLOB, MATCHED, DROPIN_EXISTS, and LOADED_NVIDIA; for GLOB, write the glob written in the configuration as it is, and for MATCHED, the number of files that glob actually catches. You check with ls <글롭> (the placeholder is the glob).
Rewrite it in containerd 2.x's configuration version 3
In /root/gpurt/etc/config-v3.toml, write the same contents in version 3 format. It is version = 3 and the plugin name is io.containerd.cri.v1.runtime. default_runtime_name is runc, runtimes must have both runc and nvidia, and nvidia has options.BinaryName. This time the imports glob must actually catch the drop-in.
containerd 2.x uses configuration version 3, and at that point the plugin name itself changes. If you raise only the version and leave the old name (io.containerd.grpc.v1.cri) as it is, that configuration is ignored entirely, yet no error appears. It is the same kind of silent failure as the previous step. After writing it, do not check by eye; read it once with a parser. python3 -c "import tomllib,sys;print(tomllib.load(open(sys.argv[1],'rb')))" <파일> (the placeholder is the file) will do.
Cross-check the handlers the cluster requires against those loaded
Create /root/gpurt/bin/rc-crosscheck.sh. If given a dump file as an argument, it reads that, and if not, it reads the result of containerd config dump. For each handler required by the cluster's RuntimeClasses that is not in the dump, it prints a line MISSING=<핸들러> (the placeholder is the handler) and ends with exit code 1. If there is none, it prints OK and ends with 0.
The key is that the basis of judgment is the loaded configuration, not the file. A check that reads /etc/containerd/config.toml cannot in principle catch this accident. That is because the file has the desired contents written in it intact and only the loaded result differs. The runtime names in the dump can be under io.containerd.grpc.v1.cri or under io.containerd.cri.v1.runtime depending on the version, so you must look at both. On the cluster side, extract with kubectl get runtimeclass -o jsonpath='{range .items[*]}{.handler}{"\n"}{end}', and do the TOML parsing with tomllib in python3. The grader actually runs this script with two kinds of dump.
If you change the default runtime, the blast radius changes
In /root/gpurt/etc/conf.d/50-default.toml, create a drop-in that sets the CRI plugin's default_runtime_name to nvidia. Write the result in /root/gpurt/out/blast.txt as four lines. They are DEFAULT_RUNTIME, AFFECTED, NEEDS_RUNTIMECLASS, and BLAST_RADIUS.
It is common for the toolkit to change the default runtime outright to nvidia. It is convenient because you need not write runtimeClassName on every Pod, but even Pods that do not use the GPU go through that runtime. If that configuration breaks, it is not the GPU Pods but all Pods on that node that cannot come up. For the values of the four lines, write in a single word each the runtime name, the scope of affected Pods, whether each Pod needs a RuntimeClass, and the scope the accident reaches.
Run the checker you made on an actual dump
In /root/gpurt/fixtures/dump-node1.toml, write an imitation of the dump of a node where only some handlers are loaded. Then run the step 6 script on that file and save the output to /root/gpurt/out/crosscheck.txt.
This fixture imitates a node where "everything is written in the file but only some of it is loaded". runc must be in it, and some of the handlers required by the cluster's RuntimeClasses must be missing. If all are present, there is no mismatch and the point of this step is lost. The grader cross-checks this fixture and the cluster by itself to compute which handlers are missing, and then matches that against your output.