TT Lab
Get started
Learn Learning paths Courses

CAPA — Argo Project Associate

What a Job Cannot Do Is What Created Workflows

Continue in TT Lab

In one line

Argo Workflows is an engine that runs several containers with ordering and dependencies. You need to start from the Kubernetes Job and understand why it falls short, and then the fields of a Workflow make sense. Events is a separate project that handles what starts that workflow.

Why this was needed

A Job expresses "run this container until it succeeds." Up to here it is perfect. The problem comes when a second task appears. If the tests must run after the build finishes, the deployment must run after the tests finish, and a notification must go out if the deployment fails, how do you link two or three Jobs? There are only three ways. Put everything into one container with a shell script, have an external orchestrator (Jenkins and so on) create the Jobs in order, or introduce a higher-level concept that can express dependencies.

The first loses the point of failure. If it dies in the middle of the script, you have to dig through the logs to find where, and a retry always starts from the beginning. The second keeps state outside the cluster, so if the orchestrator dies the pipeline is orphaned. Workflows chose the third. It put the dependencies between tasks and the passing of data into Kubernetes objects as declarations.

How it works

A Workflow's spec has a templates array and one entrypoint. The entrypoint is a name that points to which template to start from. There are several kinds of templates; the typical ones are the container template that actually runs a container, the steps template that calls other templates in order, the dag template that calls them with a dependency graph, and the suspend template that waits for human approval.

The difference between steps and dag is expressive power. steps is a list of lists, a simple model where things on the same layer run in parallel and the next layer comes after. dag lets each task write its own prerequisites directly with dependencies, so you can draw a graph such as "C runs after both A and B finish, and D runs once only B finishes." Once a pipeline goes beyond five stages, steps makes it hard to read "why this stage is here," so people usually move to dag.

Data flows along two branches. Parameters are strings and artifacts are files. If you write a path in one template's outputs.artifacts, it uploads that path to the artifact store, and the next template's inputs.artifacts downloads it and unpacks it at the specified path. The need for this intermediate store is the decisive difference from a Job. A Job has no way for Pods to pass files to each other, so it has to share a PVC or upload somewhere itself.

Reuse is handled by WorkflowTemplate. You put a frequently used template in a namespace once, and the tasks of each Workflow point to the name and the template with templateRef. It is a structure in which fixing the organization's standard build procedure in one place is reflected in every pipeline. A CronWorkflow wraps the whole Workflow spec under workflowSpec and adds schedule and concurrencyPolicy. If you set concurrencyPolicy to Forbid, a new run is skipped when the previous run has not yet finished.

Events consists of three objects: EventSource → Sensor → Trigger. The EventSource subscribes to the outside world, such as webhooks, S3, Kafka, and calendars, and sends it to the internal event bus; the Sensor subscribes to that bus and, when a condition is met, the Trigger creates a Workflow or another Kubernetes object. Why it is a separate project is answered just by looking at this list. If you start putting these adapters inside the deployment controller, Argo CD becomes not a deployment tool but integration middleware.

What it looks like in the field

The case where this distinction actually became a problem in the author's homelab was KubeVirt. The component states were all AllComponentsReady, yet the VM did not come up, and when I took apart the virt-launcher Pod spec, the volume mount holding the binary that the init container was to run was missing. The lesson confirmed a third time on this cluster alone is that "a state of Ready and actually working are different propositions."

There is exactly the same trap in workflows. A dag task ending as Succeeded only means the container exited with 0; it does not mean it actually left the artifact the next task expects. If there is no file at the path written in outputs.artifacts, that task succeeds but the next task cannot find its input and fails. So a task that hands over artifacts needs the habit of checking at the end, by itself, "whether the file is at that path." Likewise, when running GPU workloads in a workflow, if you request only nvidia.com/gpu: 1, training that needs a 32GB card can land on an 8GB laptop GPU. From Kubernetes' point of view both are one GPU, so you have to attach weight-class labels to nodes and have workloads pick with nodeSelector.

What you will do in the next lab

In /root/capa-wf/ you write a Workflow (dag dependencies, parameters, artifacts), a CronWorkflow, and a WorkflowTemplate reference, then put a Kubernetes Job and CronJob that do the same thing into a real cluster and compare the two side by side. At the end, you extract values from both the files and the cluster and build a summary ConfigMap.