TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

Where Do You Put the Gate

Continue in TT Lab

One-line summary

The gates in a deployment pipeline must be applied to the rendered result, not the source. What enters the cluster is not a file but a rendered manifest, and you cannot block what the gate does not look at.

Why the rendered result

In a repository with three kustomize overlays, you often see a gate like this.

grep -r "image:.*:latest" .   # latest 태그 금지

This check looks at the base and passes. But what is actually deployed is the result after the overlay overrides the image, and that result may contain latest. The reverse also holds. Even if the base has latest, if the overlay overrides it with a digest, what gets deployed is safe, yet the gate turns red.

Both cases happen because the gate is not looking at the thing it is trying to block. So there is only one order.

kustomize build overlays/prod   →   렌더 결과
        ↓
스키마 검증 · 정책 검사          →   이 결과를 검사한다
        ↓
0 이 아닌 종료 코드면 배포 중단

How it works

A gate must speak through an exit code

There are a great many pipelines that run a policy tool, print the result only to the screen, and stop there. The log has red text and the pipeline is green. The core of a gate is not the check but turning the check result into an exit code.

if ! kyverno apply policy/require-pinned.yaml --resource "$RENDER"; then
  echo "정책 위반으로 배포를 막습니다"; exit 3
fi

And once you have built a gate, you must actually feed it something it should block and confirm that it turns red. A gate whose only verified behavior is passing will go green for months while protecting nothing.

A digest, not a tag

checkout:1.4.0 is a name, and checkout@sha256:... is content. A tag can be pushed again, so what you verified yesterday and what is deployed today can differ. Pinning by digest lets you say that "the thing that passed the gate" and "the thing that entered the cluster" are the same. It is the minimum condition for an audit trail to hold.

In progressive delivery, the pause is the core

If you write a canary as setWeight: 10 → setWeight: 100, it is not a canary but a slightly slower full rollout. Only when there is a pause between weights is there a place for a person or an analysis to make a judgment. And because you need to be able to look at the canary by itself to have grounds for judgment, you separate canaryService and stableService.

Blue/green is an option where you check the new version through a preview path and then switch traffic. autoPromotionEnabled: false prevents automatic promotion. But the preview version can also write data through startup jobs or background jobs. For a service that needs a single writer, such as a ledger, you must verify write control, schema compatibility, and the recovery procedure separately. It does not mean that blue/green prevents concurrent writes.

What it looks like in the field

At one organization, canary analysis kept passing, yet an outage occurred. The cause was a typo in the analysis template name. The Rollout referenced a template that did not exist, and that fact came to light only right before promotion. That is why "does the referenced template actually exist?" is an item to check before deployment.

Another thing you often see is a mismatch between the destination namespace of an Argo CD Application and the namespace in which the manifests are rendered. The sync succeeds and goes green, yet nothing is reflected anywhere.

Where the green light lies

The two cases we saw earlier (a missing analysis template and a mismatched destination namespace) have something in common. The sync succeeded, but nothing happened. In GitOps, what a green light means is "git and the cluster are the same," not "the service is healthy." If you confuse the two, you feel reassured just by looking at the screen.

The places where a green light lies are few and well defined.

Nothing was created, yet it says they are the same. If the rendered result of the manifests is empty, it becomes "there is nothing to manage, so there is no difference," and the sync succeeds. This happens when you wrote the overlay path wrong or the selector selects nothing. Looking at the count of managed objects alongside it is the cheapest way to catch this.

It was applied, but the controller rejected it. The object was created, so git and the cluster are the same, but the controller that reads that object rejects the value and writes the error only in its status. Things handled by another controller, such as Rollout, Certificate, and ExternalSecret, fall under this. So you need a setting that judges health by the status field, not by the existence of the object.

An ignore rule covers real differences too. You add an ignore rule so that autoscaling changing the replica count is not treated as a difference, but if its scope is wide, it also covers things a person changed by hand. From then on, that field is outside GitOps.

So you place one more signal next to the green light. It is which commit was last synced, and what the tag of the image currently running is. If the two differ from what you expect, something is wrong regardless of the green light, and this comparison takes only a few seconds.

What to do in the next lab

You create the base and the production overlay, apply a policy gate to the rendered result, confirm that the gate actually blocks violations, and then decide which Service to attach the canary and blue/green to.