TT Lab
Get started
Learn Learning paths Courses

Kubernetes Distributions — Build Them Yourself

Why k3s reinstalls the traefik you deleted

Continue in TT Lab

One-line summary

k3s applies the server's manifests directory at startup and whenever a file changes, and it rewrites the default component files at every startup, so to remove a component you must use --disable rather than kubectl, and to change a value you must use a HelmChartConfig rather than editing the file.

Why this was needed

In managed Kubernetes, there is a separate "person who installs" the Ingress controller or the metrics server. If you don't need one, you delete it with kubectl delete, and if you need it again, you install it again. k3s was built to give you, from a one-line install, a cluster equipped with coredns, traefik, local-storage, and even metrics-server. To do that, k3s itself has to be the installer of those components, and it has to restore them to their original form even if someone breaks them by hand. So k3s keeps the manifests of the default components on the server disk and rewrites and applies them at every startup. The official documentation says that these files are rewritten at every startup for integrity, so you should not edit them.

What this design looks like in the field was confirmed by measurement (v1.35.8+k3s1).

kubectl -n kube-system delete helmchart traefik   → helm-delete Job 이 릴리스를 지움, traefik 사라짐
(20초 기다려도 그대로)
systemctl restart k3s                              → API 5초, HelmChart 가 새 UID 로 생기고 20초 뒤 traefik 복귀

The "I deleted it because I don't use it, but it's back after a reboot" incident is exactly this.

How it works

There are two layers.

The first layer is the deploy controller and AddOn. One file under /var/lib/rancher/k3s/server/manifests (including subdirectories) becomes one AddOn object in the kube-system namespace, and its name is the file name without the extension. The controller applies the file in a way similar to kubectl apply and records a content hash in the AddOn's spec.checksum. There are only two triggers for applying — k3s startup, and a change in the file content. In measurement, deleting the object from the cluster, or merely running touch on the file, did not cause it to be applied again, and when a value was changed, it was recreated with a new UID within a few seconds. Conversely, even if you delete the file, the cluster object and the AddOn remain. The documentation also states clearly that deleting a file does not delete the resource.

The second layer is the helm-controller and HelmChart. traefik.yaml contains a HelmChart object, not a Deployment, and the helm-controller sees it and starts a helm-install-traefik Job to install the chart. The chart file is fetched from the API server's static path. The precedence of values is chart defaults → HelmChart valuesContent → HelmChartConfig valuesContent → HelmChart spec.set, with the later one winning.

apiVersion: helm.cattle.io/v1
kind: HelmChartConfig
metadata:
  name: traefik          # 대상 HelmChart 와 이름·네임스페이스가 같아야 한다
  namespace: kube-system
spec:
  valuesContent: |-
    logs:
      general:
        level: DEBUG

A HelmChartConfig is a separately placed object, not a file that k3s rewrites, so it is not overwritten at startup. When you apply it, a helm upgrade runs and the release Secret count increases by one.

There are also two ways to turn components off, and they differ in nature.

--disable=traefik     AddOn 을 실제로 제거하고 원본 파일도 지운다. 구성은 기동할 때 읽힌다
traefik.yaml.skip     그 파일을 없는 것처럼 무시한다. 이미 만든 객체는 지우지 않는다

The configuration is read from /etc/rancher/k3s/config.yaml and config.yaml.d/*.yaml (in name order); for the same key the last value wins, and adding + to the end of a key appends to the list. If the file and the CLI arguments have the same key, the CLI wins. INSTALL_K3S_EXEC passed to the install script is stored as arguments of the systemd unit, so it counts as the CLI side.

What it looks like in the field

The most common case is a team trying to switch the Ingress to nginx deleting traefik with kubectl. It is quiet that day, but on the day of a node reboot or a k3s upgrade traefik comes back to life and takes ports 80 and 443, and the LoadBalancer of the new Ingress controller gets stuck in Pending. If you don't know the cause, you start by looking for "who installed it."

The second is changing values by editing traefik.yaml directly. It takes effect right away, but it reverts to the original at the next startup. Only if you use a HelmChartConfig for values and --disable for removal will the change survive reboots and upgrades.

The third is when you edited a configuration file but it was not applied. In measurement, even after writing write-kubeconfig-mode: "0600" in config.yaml and restarting, k3s.yaml remained at 644. This was because --write-kubeconfig-mode 644 passed at install time was still left as a unit argument. Writing disable: in a drop-in without + overwrites the list from the earlier file wholesale, so components that had been turned off are installed again — this is a mistake of the same family.

What really matters in practice

What you will do in the next lab

In k3s inside the VM, you record the list of AddOns, then compare deleting a user AddOn with kubectl against editing its file. You delete the HelmChart traefik, restart, and confirm by UID that it comes back to life, then apply .skip, drop-in disable, CLI precedence, and HelmChartConfig in turn and summarize in a table what requires a restart.

Reference documents: Managing Packaged Components · Helm (HelmChart·HelmChartConfig) · Configuration Options