Blue-Green or Canary — What Are You Trading
One-line summary
Blue-green and canary are not a matter of good versus bad but of what to buy among rollback speed, traffic control precision and resource cost.
Why this is needed
If you swap out the deployment all at once, then when a problem occurs, the time to roll back is the outage time itself. So deployment strategies are all devices for adjusting "how gradually to expose the new version" and "how quickly to roll back if a problem shows".
Blue-green keeps two environments of the same size up and switches the traffic all at once. Before the switch, it runs smoke tests with prePromotionAnalysis to confirm that green is healthy, and after the switch it sets scaleDownDelaySeconds to 300–600 seconds to keep the old version alive for a while. After that time passes, the old version's Pods are terminated, so a rollback has to create new Pods and takes longer. In other words, this delay is the validity period of an instant rollback.
Canary raises the percentage like 5% → analysis → 20% → 50% → 100% and judges by metrics at each step. Here the order is decisive. setWeight must run first and traffic must flow to the canary before the analysis starts, so that metrics based on real traffic accumulate. If you reverse the order, you end up judging with no data at all.
How it works
In real use, the automatic abort thresholds look like this. Success rate of 0.99 or more (interval 30s, count 10, failureLimit 2), p95 latency of 300 ms or less, error log rate of 0.005 or less. If the organizational standard SLO is a success rate of 99% and P95 of 300 ms or less, the canary thresholds should be matched to that for things to be consistent.
The most common trap is Inconclusive. If the canary traffic is too small and the query returns NaN, the verdict stays undecided. There are two responses. Give a default value like default(result[0], 1), or put a pause before the analysis to gather samples. The pipeline itself must know the fact that with 0 samples it cannot judge. So a well-made analysis script gives a failure, not a pass, when the request count is 0.
The selection criteria are summarized in a table. On traffic control precision, canary is high (1% units) and blue-green has none. On rollback speed, blue-green is very fast. On resource cost, blue-green needs twice the Pods. On data schema changes, blue-green is relatively simple. The general rule is this. Canary for most services, and blue-green for services where even a single error is costly, such as payments or authentication. If you use a rolling update, the standard values are maxSurge 1 and maxUnavailable 0.
What it looks like in the field
However well you choose the strategy, rollback breaks once a DB migration is involved. This is because, whether rolling or blue-green, there is always a window of tens of seconds to several minutes in which the old and new code are up at the same time. If you change the schema first, the old version's code breaks, and if you change the code first, the new code breaks by referencing a schema that does not exist yet. Whichever you do first, one side breaks.
The solution is the three-step Expand-Contract. In Expand you let the old and new structures coexist, in Migrate you do backfill and double writes, and in Contract you remove the old structure. The key is the asymmetry of rollback. An Expand rollback can just be ignored, and a Migrate rollback is also safe, but a Contract rollback cannot return with a simple code rollback because the old column has already been deleted. Hence the maxim: Expand freely, Contract carefully. The most common mistake is putting Expand and Contract into the same release.
Things that cannot be undone
The worth of a deployment strategy comes from being able to undo it. But some things do not come back even when you revert the traffic. Before choosing a strategy, check this list first.
What has already been sent. Email, notifications, payment approvals, requests sent to external systems. If the new version sent them wrongly, a rollback only stops sending more, and for what has already gone out, the only way is to send one more cancellation notice. So for changes with side effects that go outside, start with an especially small canary percentage, and first look at whether the point of no return can be postponed until after the deployment.
Deleted data. The Contract of the Expand-Contract we saw earlier is this. After you delete a column, even if you roll the code back, that column does not come back.
What has piled up in a changed format. If the new version wrote into a cache or queue in a new format and you roll back, the old version fails reading it. Changing the cache key together when the format changes is the way to eliminate this problem entirely, and for a queue, the old version must be able to skip the new format.
A migration that goes in only one direction. If you move data to a new store and stop using the original, from that point a rollback means losing what has piled up in the meantime.
So when making a deployment plan, you write down two things together. The procedure for undoing and when the point of no return is. If you do not write down the latter, you make the judgment "is it OK to roll back now?" for the first time in the middle of an outage. At that moment, nobody is calm.
What you will do in the next lab
This Pod has neither Argo Rollouts nor a load balancer. So you bring up two containers on 8091 (blue) and 8092 (green), express the active target with a single file, and build the switch, health check and automatic rollback in shell. You also get a feel for canary percentage distribution and the automatic abort verdict in the same way. Once you see that deciding where to send traffic is a single line in a file, what a real controller does becomes much clearer.