TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

How containerd Configuration Is Merged

Continue in TT Lab

In one line

That something was written in a configuration file and that the configuration was loaded are different events. Between them are four gaps: the merge rules, whether it was restarted, the cleanup logic when a Pod terminates, and the containers that are already running, and GPU accidents happen almost entirely in those gaps.

Why this was needed

This really happened. We edited the ClusterPolicy to turn on time-slicing, and as a consequence the toolkit DaemonSet ran again. And from that day, nobody could use the four GPUs of the cluster.

It took several hours to find the cause, and most of that time was spent checking the wrong place with precision. If you run cat /etc/containerd/conf.d/99-nvidia.toml, the nvidia runtime is written there, intact. The file is there. So you judge "it is configured" and dig elsewhere. The answer that came out when we finally typed containerd config dump was a single line — the only runtime handler loaded was runc.

How it works

Gap 1 — a drop-in can be ignored entirely. containerd's main configuration can pull in other files with imports.

version = 2
imports = ["/etc/containerd/conf.d/*.toml"]

There is a point here where almost everyone gets it wrong. This merge is not per field. If a drop-in touches the configuration of a plugin, the entire configuration of that plugin is replaced with the contents of that file. Items the drop-in did not write do not revert to the values of the earlier file; they become defaults.

Drop-ins are read in name order, so in the end the last file that mentions that plugin takes everything.

These are values actually measured on the accident node (containerd 1.7.27).

99-nvidia.toml 만 import              → runtimes.nvidia  6개
zz-labhub-registry.toml 만 import     → runtimes.nvidia  0개
둘 다 import                          → runtimes.nvidia  0개   ← 여기

zz-labhub-registry.toml was three lines. It was put in to point at the certificate path because Harbor was on HTTP.

[plugins."io.containerd.grpc.v1.cri".registry]
  config_path = "/etc/containerd/certs.d"

There is not one character about runtimes in it. Yet just because it mentioned the CRI plugin, and just because its name came after 99-, it replaced the entire CRI configuration with these three lines. The three nvidia runtimes disappeared then.

Why it stayed invisible for a week

This is the truly frightening part of this failure. If you open config dump, it comes out like this.

[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc]
  runtime_type = "io.containerd.runc.v2"
sandbox_image = "registry.k8s.io/pause:3.8"

runc is there and sandbox_image looks plausible. The configuration looks alive. In fact, everything is default. Not a single value we wrote remains, but the defaults resemble values we might have written, so it does not show.

So on the basis of this dump, we misdiagnosed twice that "the drop-in merge doesn't work". Both times it was plausible and both times it was wrong. What gave the answer was not inference but removing drop-ins one at a time and taking the dump again.

The principle of not drawing conclusions by reading applies not only to files. When you read a dump too, you must distinguish whether the value you are looking at is a value I wrote or a default. The only way to tell the two apart is to remove things one at a time and see whether the result changes.

Gap 2 — SIGHUP does not register runtime handlers. The toolkit container writes the configuration and then sends SIGHUP to containerd. But containerd 1.7's reload does not reinitialize plugins. The CRI plugin reads the runtime handler list at initialization time, so a newly added handler does not get into the list until the process is restarted. So the state "the toolkit Pod is Ready and the configuration file is updated, but the runtime is not there" is produced quite normally.

Gap 3 — if you delete the toolkit Pod, the configuration reverts. The toolkit container has cleanup logic on exit. It removes the nvidia configuration it put in and restores things to the original. By design it is right — if you remove the Operator, junk configuration must not remain on the node. The problem is the reflex during incident response. If you first do "let's delete the Pod and bring it back" and then restart containerd, the configuration is already gone by the time of the restart. If the order is reversed, it does not get better however many times you repeat.

Gap 4 — already running containers tell you nothing. The runtime handler is looked up only when a container is created. GPU Pods that are already running keep running fine even if the configuration disappears. So the failure is not visible immediately, and it bursts all at once at the moment the node is rebooted or the Pod is recreated — usually days later, usually at dawn. This latency is what makes GPU configuration accidents especially expensive.

What it looks like in the field

First, the recovery order is fixed. The order of rolling back is always this. ① Check the currently loaded handlers with containerd config dump → ② return the toolkit DaemonSet to a healthy state and have the Pod write the configuration again → ③ then restart containerd → ④ check again with config dump. If you swap ② and ③, it is back to square one.

Second, automate a post-boot check on nodes. A latent failure cannot be caught by a person's memory. One script is enough that, every time a node comes up, checks the needed handlers in containerd config dump and fails if they are missing. The key point is that a check that reads the file can never catch this accident.

Third, write the ownership of the configuration in the documentation. If several parties edit a node's containerd configuration (OS image build, configuration management tools, the GPU Operator), someday they will overwrite one another. Deciding in one line who owns which file is the only way to fundamentally reduce this accident.

What you will do in the next lab

In the next lab, you create these gaps by hand. You see that a Pod gets scheduled even with the handler of an unvalidated RuntimeClass, create a situation where an imports glob is off so that a drop-in is ignored entirely and another where the configuration version rises to 3 and the plugin name changes, and then create yourself a checker that cross-checks the handlers the cluster requires against the handlers loaded. The habit of looking at what is loaded rather than the file is contained in that one script.