Authoring and Shipping Helm Charts
Where Capabilities Come From, and Why crds Lives Outside the Release
Summary in one line
The values in .Capabilities come from a different source depending on the command, and crds/ is neither a template nor part of the release manifest but has a lifecycle of its own.
Why this is needed
Once one chart starts supporting several clusters, code like this soon comes in.
{{- if .Capabilities.APIVersions.Has "policy/v1/PodDisruptionBudget" }}
apiVersion: policy/v1
{{- else }}
apiVersion: policy/v1beta1
{{- end }}
It is a branch saying old clusters do not have policy/v1, so use the beta version. It seems to work, but when you run a render check with helm template in CI, only the else side ever comes out. In an actual deployment, on the other hand, the if side comes out. The same chart and the same values give different results.
The reason is that the source of the capability information differs.
| Command | Does it ask the cluster? | KubeVersion |
|---|---|---|
helm template |
No | The default built into Helm |
helm template --validate |
Yes | The actual cluster |
helm install --dry-run |
Yes | The actual cluster |
helm install --dry-run=server |
Yes | The actual cluster |
helm install |
Yes | The actual cluster |
Only helm template is offline. And its default capability list is very small, so even networking.k8s.io/v1/Ingress, which an actual cluster naturally has, comes out as missing. If you do not know this, you start fixing the chart, saying "I rendered it and the Ingress branch is not taken."
To test branches without a cluster, you just have to make Helm lie. --kube-version 1.21.0 adds a version, and --api-versions "demo.labhub.io/v1/Widget" adds to the API list. The fact that it adds is important — it does not replace the default list but adds to it.
The crds directory differs in five ways
CRDs have a chicken-and-egg problem. For the same chart to also create the custom resources a CRD defines, the CRD must already be in the cluster first. Helm solves this with a dedicated directory called crds/. That directory differs from ordinary templates in five ways.
- It does not go through the template engine. Braces are just characters even if you use them. So you cannot branch a CRD conditionally on values.
- It does not appear in the default render result. It comes out only if you give
helm template --include-crds. - It goes in before everything else at install time. So the same chart's templates can create objects that use that CRD together, and
.Capabilities.APIVersions.Hasbecomes true even on the first install. - It is not touched on upgrade. Even if you edit the chart's
crds/and runhelm upgrade, the CRD in the cluster stays as it is. The official documentation states this as a limitation and tells you to update CRDs by hand. - It remains even when you delete the release. Deleting a CRD would make all the custom resources created from it disappear with it, so Helm hands this decision to a person.
If you pull out the release manifest with helm get manifest, the CRD is not there. It means the release does not own the CRD, and the fourth and fifth points above follow from this.
An alternative — putting the CRD in templates
If the constraints of crds/ (not updated on upgrade, no conditional branching) are a problem, there is the option of putting the CRD in templates/. Then it becomes an ordinary object that is updated by upgrade and can have conditions attached. In exchange, the release comes to own the CRD, so deleting the release makes the CRD and all its resources disappear. And if several releases use the same CRD, ownership overlaps and they conflict.
The boundary often used in practice is this. Make the CRD a separate chart split off from the operator chart, have the cluster administrator install it only once, and have the application chart only check for its existence with .Capabilities.APIVersions.Has. This is why large projects publish a separate <이름>-crds chart (with the project name in the placeholder).
What it looks like in the field
The most frequent accident is "I fixed the CRD and upgraded, but the new field has no effect." Revisions pile up in helm history and the deployment is shown as a success, yet the CRD in the cluster is still the old schema. A custom resource that uses the new field has that field treated as not being in the schema and quietly trimmed away. What makes this combination especially bad is that no failure is reported anywhere. The standard response is to put a step in the deployment pipeline that applies the CRD separately with kubectl apply, or to keep a dedicated chart just for CRDs.
The second most common is the CI render check producing a different result from production. If you only run helm template, the capability branches are always fixed to one side. If CI can use a real cluster, use --validate or --dry-run=server, and if it cannot, it is safer to imitate the target clusters with --kube-version and --api-versions and render both cases.
What you will do in the next lab
You build a chart that exports the capabilities as they are and check the values of an offline render, then render pretending to be a different cluster with --kube-version and --api-versions. You create a branch that puts in or leaves out a custom resource depending on capabilities, place a CRD in crds/, install it on a real kwok cluster, and confirm that the CRD is not in the manifest. Finally, you see for yourself that the cluster stays as it is even if you fix the CRD and upgrade, and sort out how --validate and --dry-run=server differ from an offline render.