How Does kubectl scale Know About My Type?
In one sentence
kubectl scale and the HPA work without knowing the kind of their target. All they know is a single shared window, the scale subresource, and a CRD opens that window with three lines: specReplicasPath, statusReplicasPath, and labelSelectorPath.
Why this contract was needed
Kubernetes has several types that have the concept of a "replica count": Deployment, ReplicaSet, StatefulSet, ReplicationController, and a great many CRDs. Their field names can differ, and in practice they do. Yet kubectl scale is a single command and there is just one HPA controller. If these tools had different code for each type, you would have to change kubectl every time a new CRD appeared. So Kubernetes went the other way — the tools are fixed, and each type translates itself and offers that translation.
The result of that translation is the Scale object of autoscaling/v1. When any type enables the scale subresource, GET .../shards/a/scale returns a Scale rather than a Shard. Inside are just one spec.replicas field, one status.replicas field, and one status.selector string. Both kubectl and the HPA look only at these three fields.
How it works
For each version of the CRD, you write three paths under subresources.scale.
subresources:
status: {}
scale:
specReplicasPath: .spec.replicas
statusReplicasPath: .status.replicas
labelSelectorPath: .status.selector
| Path | Who uses it | If missing or wrong |
|---|---|---|
specReplicasPath |
Users, kubectl, and the HPA write it | Required. It must be under .spec. If the value is not in the actual object, the subresource returns an error |
statusReplicasPath |
The controller writes it to status | Required. It must be under .status. If the object has no value, the actual count looks like 0 |
labelSelectorPath |
The controller writes it to status and the HPA reads it | Optional. Without it, the HPA cannot start adjusting |
All three paths allow only dot notation, not array notation. The place labelSelectorPath points to must be a single string, not a structure, and it holds the label selector serialized (app=web).
There is something you must remember together with this. The place statusReplicasPath points to is always inside the status subresource, and labelSelectorPath is usually placed there too (under .spec is also allowed, but the selector is a value the controller computes and fills in, so status is where it belongs). So if you do not also enable status: {}, the controller has no window to write that value to, and even if you enable it, an ordinary kubectl patch will not write it — you need --subresource=status. If you run kubectl scale on a cluster with no controller, only the desired count changes to 5 and the actual count stays at 0, which is not a malfunction but means that nobody has yet filled in that place.
The HPA side is a bit more interesting. The HPA finds its target by the apiVersion, kind, and name written in scaleTargetRef, and reads that target's scale subresource. Whether this succeeded is recorded in AbleToScale of status.conditions, and whether it can actually adjust after computing the metrics is recorded in ScalingActive. These are two different questions. If the selector path is missing, AbleToScale is True (SucceededGetScale) but ScalingActive is False (InvalidSelector), and the message says that the target's scale has no selector. In the list view, the HPA still shows up as a perfectly normal single line.
What you see in the field
First, a failure that answers success. If you write specReplicasPath with a one-character mistake, such as .spec.replica, kubectl scale prints scaled with exit code 0. But the object's value is unchanged. Only when you ask the same CRD with kubectl get … --subresource=scale does "the spec replicas field ... does not exist" appear. If a deployment script just looks at the exit code and moves on to the next stage, this defect lives for months.
Second, an HPA that is attached but does not adjust. If you attach an HPA to a CRD that left out the selector path, nobody sees an error. Even as the load rises, the replicas stay the same. People usually suspect the metrics pipeline first, but the real cause is that one of the CRD's three lines is missing. The reason it takes so long to reach this cause in an incident review is that the symptom is "autoscaling isn't working," so the investigation opens up toward the metrics side first.
Third, so you automate the check. A team that builds CRDs does well to verify two things in CI — actually reading the scale subresource once, and, if the type is to be autoscaled, checking that a selector path exists. You cannot learn the first by reading the definition alone. Whether the path actually reaches the real object comes out only when you ask.
Limits of this lab environment
The cluster in the lab Pod has no metrics server (metrics-server). So the HPA's TARGETS column stays unknown, and you cannot see the replica count actually rise and fall with load. Instead, the HPA controller itself does run, so "did it read the target" and "why could it not start adjusting" are recorded in the conditions as they are — every failure this module deals with is left in that place. Also, since no controller looks after Shard, you fill in status by hand to imitate the controller's place.
What you will do in the next lab
You create a CRD with all three paths and attach kubectl scale and an HPA to it, then create a CRD without the selector path and a CRD with a typo in a path, and check from the condition messages what kind of silence each produces. Finally, you build a check script that groups the three CRDs in a table and classifies them as ok, nosel, or broken, and leave the output from before and after the fix side by side.