TT Lab
Get started
Learn Learning paths Courses

CAPA — Argo Project Associate

Between Declaration and Behaviour

Continue in TT Lab

In one line

Argo's objects only record intent, and what turns that intent into reality is the controller. Without a controller, you cannot even tell whether the YAML is right.

Why a controller is needed

In the earlier modules you wrote several Workflows and Rollouts. But the place where those labs ran had no controller, so there was no way to check what that YAML actually does.

All four Argo projects are things that only make sense with a controller. Objects only record intent, and what turns that intent into reality is the controller.

Workflows   컨테이너를 실제로 돌린다      없으면 단계가 진행되지도 실패하지도 않는다
Rollouts    파드를 실제로 띄우고 나눈다    없으면 카나리가 몇 퍼센트인지 볼 수 없다
Analysis    Job 을 띄워 판정한다          없으면 자동 롤백이 통째로 빠진다

The Korean text in this code block lists three controllers in order and what is missing without each: Workflows actually runs containers, and without it steps neither progress nor fail; Rollouts actually starts and splits Pods, and without it you cannot see what percentage the canary is; Analysis starts a Job to judge, and without it automatic rollback is missing entirely.

What went back is the traffic, not the declaration

This is the most valuable trap. Even if the analysis fails and an automatic rollback happens, the image in spec.template stays at the new version as it is.

To fix it, you have to revert the declaration — a commit that reverts the image tag in the repository.

limit is the retry count

The retryStrategy.limit of Workflows is the number of retries. The first attempt is not counted in it. With limit: 2, the actual attempts are 3.

Fixing the template mid-run does not change it

The controller embeds the definition at start time into status.storedTemplates. This is for reproducibility. This is the answer to "I fixed the template, so why is it the same?", and the change takes effect from the next workflow.

What a canary weight actually does

Whether setWeight: 20 means "20% of requests" or "20% of Pods" differs depending on the configuration. This difference makes measured values and expectations diverge.

trafficRouting Meaning of the weight Accuracy
None Ratio of Pod count Coarse. With replicas 5, 20% is 1 Pod
Istio·Gateway API Actual request ratio Accurate
NGINX Ingress Actual request ratio (annotation) Accurate

When there is no trafficRouting, the controller approximates with the Pod count. If you give setWeight: 10 at replicas: 3, you cannot create 0 Pods, so it becomes 1 (about 33%). To start with a small weight, replicas must be large enough or there must be a trafficRouting.

What the analysis looks at to judge

An AnalysisTemplate queries a metric provider and decides whether to continue from the result.

metrics:
  - name: error-rate
    interval: 1m
    count: 5                    # 5번 재고
    failureLimit: 2             # 2번 실패하면 롤백
    successCondition: result[0] < 0.05
    provider:
      prometheus:
        address: http://prometheus.monitoring.svc:9090
        query: |
          sum(rate(http_requests_total{status=~"5..",version="{{args.version}}"}[2m]))
            / sum(rate(http_requests_total{version="{{args.version}}"}[2m]))

What people often get wrong here is a query that does not filter for only the canary. If the above query has no version label, it measures the stable version and the canary combined, so even if the canary fails 100%, the overall error rate rises only 20%. It does not exceed the threshold, so a bad deployment passes.

And there is a trap when traffic is low. If 3 requests reach the canary and 1 fails, the error rate is 33%. You must also set a minimum sample count condition.

successCondition: result[0] < 0.05 || result[1] < 20   # 요청이 20건 미만이면 통과

Criteria for choosing a deployment strategy

Canary Blue-green
Resources A little more (per step) Double
Rollback speed Set the weight to 0 — fast Switch the Service selector — fastest
DB schema The two versions coexist for a long time The two versions coexist briefly
Suitable when Traffic is heavy and metrics can be used to judge Verification is done by a person or traffic is low

If you use a canary on an internal service with little traffic, the sample is too small and the analysis is meaningless. In that case it is better to use blue-green and have a person check the preview before moving on.

What really matters in practice

After an automatic rollback, always add a commit that reverts the declaration. A rollback only reverts the traffic, and spec.template stays at the new version. If you use GitOps, the repository still points to the new image, so it retries the same deployment endlessly.

Judge a deployment not by Healthy but by status.stableRS. Even after promotion finishes and it becomes Healthy, the old Pods briefly remain as Terminating, so if you count images, you get two kinds. The hash that points to the stable version is the only criterion that does not wobble.

retryStrategy.limit is not the total number of attempts. The first attempt is not counted in it, so limit: 2 actually runs 3 times. It is a spot where everyone omits one when calculating a timeout budget.

In the next two labs, you confirm these things directly on real controllers.