TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

You raised the replicas and nothing happened - the scale subresource contract

Continue in TT Lab

Goal

Bind the CRD's subresources.scale properly with three paths, build for yourself the silent failure that occurs when you leave out the selector path and when there is a typo in a path, and then create a check script that catches those silent failures.

Why it matters

kubectl scale and the HPA do not know the kind of their target. All they know is a single shared window, the scale subresource, and a CRD opens that window with three JSONPaths. So if the person building an Operator writes these three lines correctly, their type connects for free to every existing Kubernetes autoscaling tool, and if they write them wrongly, those tools do nothing, without an error. This silence is the subject of this module. If there is a typo in a path, kubectl scale answers that it succeeded while not changing the value, and if the selector path is missing, the HPA succeeds in reading the target and then never adjusts, forever. Both stay hidden until a person looks at the screen, so you need to have a check method that you can put in a deployment pipeline.

Steps

  1. In /root/op-scale/shard-crd.yaml, write the CRD shards.scale.labhub.io — group scale.labhub.io, kind Shard, plural shards, a single version v1 (both served and storage true), put spec.replicas (integer), status.replicas (integer), and status.selector (string) in the schema, and enable status: {} and scale in subresources. The three paths of scale are .spec.replicas, .status.replicas, and .status.selector respectively. After applying, save the three paths registered in the cluster to /root/op-scale/scale-paths.txt as three lines in <이름>=<값> (name=value) form (specReplicasPath=..., statusReplicasPath=..., labelSelectorPath=...).
  2. Create the namespace op-scale and write Shard a in /root/op-scale/shard-a.yaml — spec.replicas is 2. After applying, save the output of kubectl -n op-scale get shard a --subresource=scale -o yaml as it is to /root/op-scale/scale-initial.yaml.
  3. Raise Shard a's desired count to 5 with kubectl -n op-scale scale shard a --replicas=5. Then take the two numbers from the scale subresource and save them to /root/op-scale/after-scale.txt as two lines, spec.replicas=5 and status.replicas=0.
  4. Write status in place of what the controller would do — fill status.replicas with 5 and status.selector with app=shard-a using kubectl -n op-scale patch shard a --subresource=status. Then save the scale subresource output to /root/op-scale/scale-ready.yaml.
  5. In /root/op-scale/shard-hpa.yaml, write an autoscaling/v2 HPA shard-a-hpa — scaleTargetRef has apiVersion scale.labhub.io/v1, kind Shard, and name a, with minReplicas 2, maxReplicas 10, and a metric of Resource cpu with Utilization 70. Apply it, wait until the conditions are filled in, and then save two lines to /root/op-scale/hpa-status.txt — AbleToScale=<상태>/<이유> (status/reason) and REFERENCE=<kubectl get hpa 의 REFERENCE 칸> (the value from the REFERENCE column of the kubectl get hpa output).
  6. In /root/op-scale/block-crd.yaml, write the CRD blocks.scale.labhub.io (kind Block, plural blocks), but do not put labelSelectorPath in scale. Create and apply Block b (spec.replicas 3) with /root/op-scale/block-b.yaml and the HPA block-b-hpa (minReplicas 1, maxReplicas 6, cpu 70) with /root/op-scale/block-hpa.yaml, and once the conditions are filled in, save two lines to /root/op-scale/noselector.txt — ScalingActive=<상태>/<이유> (status/reason) and MESSAGE=<조건 메시지> (the condition message).
  7. In /root/op-scale/relay-crd.yaml, write the CRD relays.scale.labhub.io (kind Relay, plural relays), but deliberately write specReplicasPath as .spec.replica (a typo with the trailing s dropped). Create and apply Relay r (spec.replicas 2) with /root/op-scale/relay-r.yaml, run kubectl -n op-scale scale relay r --replicas=4 and kubectl -n op-scale get relay r --subresource=scale -o yaml in turn, and collect their results in /root/op-scale/typo-report.txt — it must have scale-rc=<종료 코드> (the exit code), spec.replicas=<명령 뒤 실제 값> (the actual value after the command), and a line containing the error sentence of the second command exactly as it was.
  8. In /root/op-scale/scale-audit.tsv, write three lines in the form <CRD 이름><탭><기대 분류> (CRD name, a tab, the expected classification) — ok for shards.scale.labhub.io, nosel for blocks.scale.labhub.io, and ok for relays.scale.labhub.io. /root/op-scale/scale-audit.sh reads this table and classifies each CRD as ok, nosel, or broken, prints OK … if it matches and MISMATCH … if it does not, to standard output only, and must exit with a nonzero code if even one line is wrong. First run it once before fixing and leave the output in /root/op-scale/scale-audit-before.txt, then fix the typo with /root/op-scale/relay-crd-fixed.yaml and apply it, check that kubectl -n op-scale scale relay r --replicas=4 actually changes the value this time, and save the output of running it again to /root/op-scale/scale-audit.txt.

Notes

Bind the contract with three paths

In /root/op-scale/shard-crd.yaml, write the CRD shards.scale.labhub.io — group scale.labhub.io, kind Shard, plural shards, a single version v1 (both served and storage true), put spec.replicas (integer), status.replicas (integer), and status.selector (string) in the schema, and enable status: {} and scale in subresources. The three paths of scale are .spec.replicas, .status.replicas, and .status.selector respectively. After applying, save the three paths registered in the cluster to /root/op-scale/scale-paths.txt as three lines in <이름>=<값> (name=value) form (specReplicasPath=..., statusReplicasPath=..., labelSelectorPath=...).

The three paths are a contract that tells the API server "where the desired count is in this type, where the actual count is, and where the selector that picks the Pods is." A path is a JSONPath that starts with a dot, and only dot notation is allowed (array notation cannot be used). If you pull the registered values out with kubectl get crd shards.scale.labhub.io -o jsonpath, you avoid the mistake of copying them by hand.

What does the scale subresource look like

Create the namespace op-scale and write Shard a in /root/op-scale/shard-a.yaml — spec.replicas is 2. After applying, save the output of kubectl -n op-scale get shard a --subresource=scale -o yaml as it is to /root/op-scale/scale-initial.yaml.

A subresource is looking into the same object through a different window. What comes back is not a Shard but an autoscaling/v1 Scale object, which holds just one spec.replicas field and one status.replicas field. Nobody has written status yet, so try to predict what the status number will be.

What kubectl scale changes

Raise Shard a's desired count to 5 with kubectl -n op-scale scale shard a --replicas=5. Then take the two numbers from the scale subresource and save them to /root/op-scale/after-scale.txt as two lines, spec.replicas=5 and status.replicas=0.

kubectl scale is not a Deployment-only command — it works as it is on every type with the scale subresource enabled. Think about why the two numbers differ. specReplicasPath is the place users write and statusReplicasPath is the place the controller writes, and this cluster has no controller looking after Shard.

Fill in the controller's place by hand

Write status in place of what the controller would do — fill status.replicas with 5 and status.selector with app=shard-a using kubectl -n op-scale patch shard a --subresource=status. Then save the scale subresource output to /root/op-scale/scale-ready.yaml.

Status is a separate endpoint, so an ordinary patch does not write it. You must add --subresource=status to go in through that window. The selector is a single string line and uses the label selector syntax (키=값, key=value) as it is — this value later becomes the basis on which the HPA counts Pods.

Attach an HPA to a custom resource

In /root/op-scale/shard-hpa.yaml, write an autoscaling/v2 HPA shard-a-hpa — scaleTargetRef has apiVersion scale.labhub.io/v1, kind Shard, and name a, with minReplicas 2, maxReplicas 10, and a metric of Resource cpu with Utilization 70. Apply it, wait until the conditions are filled in, and then save two lines to /root/op-scale/hpa-status.txt — AbleToScale=<상태>/<이유> (status/reason) and REFERENCE=<kubectl get hpa 의 REFERENCE 칸> (the value from the REFERENCE column of the kubectl get hpa output).

The HPA controller does not need to know the kind of its target. It attaches as long as it can read the scale subresource. You can pull the conditions out with kubectl -n <ns> get hpa <이름> -o jsonpath (replace the placeholders with the namespace and the HPA name), and in this environment without a metrics server, TARGETS stays unknown — that is a different matter from "can it read the target."

Leave out the selector path and the HPA quietly stops

In /root/op-scale/block-crd.yaml, write the CRD blocks.scale.labhub.io (kind Block, plural blocks), but do not put labelSelectorPath in scale. Create and apply Block b (spec.replicas 3) with /root/op-scale/block-b.yaml and the HPA block-b-hpa (minReplicas 1, maxReplicas 6, cpu 70) with /root/op-scale/block-hpa.yaml, and once the conditions are filled in, save two lines to /root/op-scale/noselector.txt — ScalingActive=<상태>/<이유> (status/reason) and MESSAGE=<조건 메시지> (the condition message).

What differs from the previous step is the key point. Reading the target (AbleToScale) succeeds, but adjusting never starts. To compute the target utilization, the HPA must know "which Pods belong to this workload," and the place that provides that answer is exactly the path you left out. Read the condition message as it is.

It answers success while doing nothing

In /root/op-scale/relay-crd.yaml, write the CRD relays.scale.labhub.io (kind Relay, plural relays), but deliberately write specReplicasPath as .spec.replica (a typo with the trailing s dropped). Create and apply Relay r (spec.replicas 2) with /root/op-scale/relay-r.yaml, run kubectl -n op-scale scale relay r --replicas=4 and kubectl -n op-scale get relay r --subresource=scale -o yaml in turn, and collect their results in /root/op-scale/typo-report.txt — it must have scale-rc=<종료 코드> (the exit code), spec.replicas=<명령 뒤 실제 값> (the actual value after the command), and a line containing the error sentence of the second command exactly as it was.

This step is not about "what happened" but "why does it look like success when nothing happened." If you write the exit code and the actual value separately, it is obvious at a glance that the two disagree. The error of the second command goes to standard error, so capture it too with 2>&1.

Build a check table that catches the quiet failures

In /root/op-scale/scale-audit.tsv, write three lines in the form <CRD 이름><탭><기대 분류> (CRD name, a tab, the expected classification) — ok for shards.scale.labhub.io, nosel for blocks.scale.labhub.io, and ok for relays.scale.labhub.io. /root/op-scale/scale-audit.sh reads this table and classifies each CRD as ok, nosel, or broken, prints OK … if it matches and MISMATCH … if it does not, to standard output only, and must exit with a nonzero code if even one line is wrong. First run it once before fixing and leave the output in /root/op-scale/scale-audit-before.txt, then fix the typo with /root/op-scale/relay-crd-fixed.yaml and apply it, check that kubectl -n op-scale scale relay r --replicas=4 actually changes the value this time, and save the output of running it again to /root/op-scale/scale-audit.txt.

The value of the check table is proved by the output from "before the fix." If you keep only the all-OK output after the fix, nobody knows what this check can catch. Classification is not enough from the CRD definition alone — whether a path actually reaches the real object can be known only by reading the scale subresource once. If the script writes files itself, it overwrites the outputs when you rerun it, so emit only to standard output.