TT Lab
Get started
Learn Learning paths Courses

CAPA — Argo Project Associate

Why Argo Split Into Four

Continue in TT Lab

In one line

Argo is not a single deployment tool but a bundle of four CNCF projects. Argo CD solves "is the cluster in the same state as Git?", Workflows solves "run these tasks to the end in a fixed order," Events solves "when something happens outside, then do this," and Rollouts solves "how slowly should a new version be rolled out?" The reason the four were not merged into one is that the nature of the time each one deals with is different.

Why this was needed

Deployment before GitOps was mostly push-based. The CI server finishes the build and fires off kubectl apply. Three costs hide in this structure. First, the CI server has to hold the credentials of the production cluster. If one Jenkins server is breached, every cluster is breached. Second, if someone changes replicas by hand after the apply, nobody knows. Git is only a record of the moment of deployment, not evidence of the current state. Third, the moment there are several clusters, the pipelines multiply by the number of clusters.

The pull model flips all three at once. A controller that lives inside the cluster reads Git and aligns itself. Credentials never leave the cluster, drift shows up as a state because the controller keeps comparing, and there is still only one Git repository even as clusters grow. This is exactly what Argo CD does, and the other three are satellites attached around this axis.

How it works

Project Problem it solves Nature of time
Argo CD Keeps erasing the difference between the declared state and the actual state An endless loop (default 180-second period)
Argo Workflows Links several container tasks into a DAG and runs them to completion once A finite run with a start and an end
Argo Events Receives an external signal and starts something An event whose arrival time is unknown
Argo Rollouts Grows a new version gradually while watching metrics A few minutes to a few hours, with human judgment involved

Here you can see why they were not merged. Argo CD is something that runs "until the difference is 0," so it has no concept of completion. Workflows is the opposite: completion is everything. If you put the two in one controller, the retry semantics, the failure handling, and even the metric names all collide. Events exists separately for the same reason. If you start putting webhook handling into Argo CD, GitHub, S3, Kafka, and even calendars all come inside the deployment controller. Events pulled that adapter hell out into a separate axis of EventSource → Sensor → Trigger.

Argo CD is a CNCF Graduated project. The practical implication of the graduated level is "this API will not be torn up next quarter," and that becomes the basis for judging that it is acceptable to make the Application manifest an organizational standard.

What it looks like in the field

The author's homelab has 3 control plane nodes and 4 GPU workers, 7 nodes in total. MetalLB hands out 10.0.0.200 through 215 over L2, and on top of that Gitea runs at 10.0.0.200, ArgoCD at 10.0.0.201, and Harbor at 10.0.0.202. The source, the GitOps controller, and the registry sit side by side in the same address range, and the real value of this layout showed when an outage happened.

In this cluster, controlPlaneEndpoint was hard-coded to the physical IP of the first control plane node rather than to a VIP. There were 3 etcd members, so quorum was fine, but when that node died, both kubectl and kubelet became unable to connect at all. The control plane was alive, but nobody could find the door. Two things were learned here: one is that "a state of Ready and actually working are different propositions," and the other is that even while the cluster was entirely invisible, the desired state remained intact in the Git repository. With a push pipeline the recovery procedure would have been "rebuild the pipeline," but in the pull model, the moment the cluster is revived, the controller catches up on its own. Many articles explain the reason for adopting GitOps as deployment convenience, but the moment its value really shows is at a recovery time like this.

Where the four projects actually overlap

The four Argo projects can be used separately, but when used together, there are places where the boundaries blur. This is where things split, both on the exam and in the field.

CD and Rollouts collide over ownership. While Rollouts is progressing a canary, the Deployment's replicas or image change, and CD sees that as a difference from git. If you do not exclude those fields with ignoreDifferences, there is a loop in which CD reverts the canary.

Workflows and Events are divided by "what is the trigger." Workflows is an engine that runs work with steps, and Events is the layer that receives outside occurrences (webhooks, queue messages, schedules) and calls that engine. With only Workflows and no Events, a person has to start it, and with only Events and no Workflows, what you can do with a received event stays simple.

Where sync order is needed, use hooks and waves. The namespace has to exist before the CRD goes in, and the CRD has to exist before the resources that use it go in. If you give no order, the first sync fails and succeeds by chance on a retry. Write the order with argocd.argoproj.io/sync-wave, and put things that must run only once, such as migrations, in a PreSync hook.

An app that creates apps (App of Apps) is convenient but hard to undo. Whether deleting the parent app also deletes the children depends on the finalizer, and it really happens that someone deletes it without knowing this and the whole cluster is emptied. Set the deletion policy of the parent and the children explicitly.

Either way, the real difficulty is permissions. For CD to deploy to several clusters, it has to hold those clusters' permissions, which means if you breach CD, everything is breached. That is why AppProject, which restricts the target namespaces and resource kinds per project, is the default and not an option.

What to check in the next quiz

This module covers concepts only. In the module that follows immediately, you write an Application manifest yourself and check by hand which fields assemble the "loop that keeps erasing the difference" just described.