Why a Deployment's Rolling Update Is Not Enough
In one line
A Rollout is a workload resource that replaces a Deployment. There is only one reason to replace it: a Deployment's rolling update has no concept of "pause briefly and watch." A Rollout puts steps called setWeight, pause, and analysis in that place.
Why this was needed
A Deployment's RollingUpdate decides the replacement speed with two numbers, maxSurge and maxUnavailable. The premise of this model is "if the new Pod is Ready, it is fine." But most deployment incidents are not incidents where Pods fail to come up; they are incidents where the Pods came up fine but the responses were wrong. While the readinessProbe returns 200, the payment API can spew 5xx, and latency can triple. A Deployment has no eyes to see this, so it pushes the problem version all the way through.
There is also no place for a human to step in. There is kubectl rollout pause, but it is a command a person has to type by hand at the right moment, so it cannot be put into automation. So the actual practice usually runs on the manual procedure of "deploy, watch Grafana, and if something looks off, rollout undo," and the quality of this procedure depends on the concentration of whoever is on call that day. A Rollout moves this procedure into the object as a declaration.
How it works
The canary strategy is expressed with a steps array. There are three representative kinds of steps.
| Step | Meaning |
|---|---|
| setWeight | Raise the traffic ratio (or Pod ratio) sent to the new version to this much |
| pause | Stop. If you give a duration, for that long; if not, indefinitely until a person promotes |
| analysis | Start an AnalysisRun to measure metrics, and if it fails, abort and roll back the rollout |
A pause without a duration is important. It is the way to put human approval into the graph, and the standard idiom for placing a manual gate in the middle of an automated pipeline.
What the analysis step references is an AnalysisTemplate. It contains a metrics array, and each metric has an interval (how often to measure), a successCondition or failureCondition (what counts as success), a failureLimit (how many failures to tolerate), and a provider (where to fetch from). If you write the successCondition like result[0] >= 0.95 and set the provider to Prometheus, the moment the success rate drops below 95% failureLimit times, the rollout stops on its own and goes back. This is the point where a procedure that depended on human concentration became a declaration.
To actually split traffic, you need the cooperation of the data plane. A Rollout can also change only the Pod ratio (with 4 replicas and setWeight 25, that is one), but precise ratio control must be done by an Ingress controller or a service mesh. That is why under trafficRouting there are implementation-specific settings such as nginx, istio, and alb, and you specify two Services together, canaryService and stableService. The controller manipulates the selectors of the two Services to change which set of Pods sits behind which Service.
Blue-green has a different shape. It brings up the new version in full and lets it be reached only through previewService, and at the moment of promotion it swaps activeService over to the new one. If you set autoPromotionEnabled to false, it waits until a person promotes, and scaleDownDelaySeconds decides how many seconds to keep the old version alive after promotion. The reason this value is not 0 is to leave room for an immediate rollback.
What it looks like in the field
The incident in the author's homelab that intersected with this topic was the GPU Operator. The container runtime configuration was off, yet the status display was normal, and it only showed up when I actually started a Pod. The gap between the status a deployment tool shows and what the service actually does is a recurring theme in this course, and the Rollout's analysis step is a device built precisely to close that gap. It makes you look not at the signal that a Pod is Ready but at the signal of response success rate.
One more thing. What you run into most in practice when switching to a Rollout is coexistence with the existing Deployment. If a Deployment and a Rollout with the same selector exist at the same time, the two controllers each claim the same Pods as their own. So the switch is done by having the Rollout reference the existing Deployment with workloadRef, or in the order of scaling the Deployment to 0 and then bringing up the Rollout. And the foundation of rollback is still held by Kubernetes. A Deployment keeps ReplicaSets as revisions, so rollout undo can go back to the previous state, and if you use a Rollout without understanding this mechanism, rollback looks like magic.
What you will do in the next lab
In /root/capa-rollout/ you write a canary Rollout, an AnalysisTemplate, and a blue-green Rollout, put the corresponding two Services and a Deployment into a real cluster, then change the image and revert it with rollout undo to see for yourself how revisions stack up.