TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

Schema, Subresources, Versions — the Three Axes of a CRD

Continue in TT Lab

In one sentence

Writing a CRD is not about listing fields; it is about designing how much of the validation you can hand off to the API server.

Why it was needed

When checks like if spec.Replicas < 1 { return error } start piling up in controller code, two problems appear. First, those checks run too late. The bad object has already been stored in etcd, and the user believes kubectl apply succeeded. Second, the rules are documented nowhere. Users have to read the controller logs to find out what went wrong.

Moving them into the schema reverses both. kubectl apply fails on the spot, the error message tells you which field is wrong and why, and for an enumeration it even lists the allowed values. And kubectl explain webservice.spec becomes the documentation as it is.

How it works

Three pillars support a CRD.

1) Structural schema. Since Kubernetes 1.16, every CRD must declare the type of every field with OpenAPI v3. Only when this condition is met do pruning, default injection, server-side apply, and CEL validation work. Three things must be clearly distinguished here.

Category Meaning When to use
required Rejected if empty Core identifiers with no reasonable default
default The API server fills it in if empty Fields where most people use the same value
Nothing attached Allowed to be empty, not filled in Truly optional features

If you make a field that could have a default required, users end up writing the same value every time as boilerplate, and it becomes hard to demote it to optional later. Conversely, for a field like the image, where a wrong default is dangerous, explicit rejection is better.

Pruning is the default behavior. Fields not in the schema are cut off before storage. It is disorienting at first when a typo'd field quietly disappears, but this is the mechanism that enforces "the schema is the contract." Where you must accept arbitrary key-value pairs, open an exception with x-kubernetes-preserve-unknown-fields: true, but open only that spot, and keep it minimal. If you overuse it, you lose the benefits of the structural schema altogether.

2) Subresources. When you enable the status subresource, spec and status become different endpoints. Users write only spec, and the controller writes only status. One decisive property follows — writing status does not raise metadata.generation. Thanks to this, the controller can tell "the user changed the spec" apart from "I just wrote status," and this is the foundation that prevents infinite reconciliation.

The scale subresource attaches kubectl scale and the HPA to your type just by telling it three paths: specReplicasPath, statusReplicasPath, and labelSelectorPath. The selector path is needed because the HPA uses that selector to count Pods.

3) Versions. A single CRD can serve several versions at once, and each version has two flags.

Because storage happens in only one representation, all versions must be convertible to one another without loss. And status.storedVersions records "the versions that have ever been used to store objects of this CRD." If you change the storage version and do not re-store existing objects, the old version remains in this list, and if you delete the old schema in that state, the stored objects can no longer be read. Most version-removal incidents come from skipping this re-storage step.

Note that this lab environment has no way to serve a webhook endpoint, so conversion webhooks are not covered. Instead, you serve several versions together with strategy: None and check the storage version rule.

What you see in the field

First, the first day you get stuck on the CRD naming rule. metadata.name must be <복수형>.<그룹> (the plural name followed by the group). If you write only webservices, the API server rejects it. The broken CRD provided as a fixture is exactly that case. The reason for this rule is that the CRD itself is cluster-scoped and its name is the globally unique key.

Second, custom columns change operational quality. If kubectl get webservices shows only NAME and AGE, nobody uses that command. If you surface the two or three things operators want to know during incident response as columns, that one command becomes a dashboard.

Third, an enumeration is documentation. When you add an enum, the rejection message includes the allowed list. Users can find the answer just by reading the error message, without digging through a wiki. The real benefit of moving validation into the schema is not the rejection itself but this guidance.

What you will do in the next labs

Two labs follow. In the first, you write the WebService of apps.labhub.io/v1 from scratch — the naming rule, schema, custom columns, the status and scale subresources, and two versions, v1alpha1 and v1. In the second, you throw resources that deliberately break the rules at the API server and collect how it words each rejection, see default injection and pruning with your own eyes, and then put cross-field constraints into the schema with CEL rules and build a validation matrix.