TT Lab
Get started
Learn Learning paths Courses

CAPA — Argo Project Associate

DAG Pipelines and How Far They Are From a Job

Continue in TT Lab

Goal

Express Argo Workflows DAGs, artifacts, and template reuse as manifests, and also actually deploy the same work as a Kubernetes Job and CronJob, to see with your own eyes the difference in expressive power between the two models.

Why it matters

To judge whether to adopt Workflows, you need to know how far a Job can go. A Job can express retries (backoffLimit) and parallelism (parallelism), but it cannot express dependencies between tasks or file passing. So an organization that uses only Jobs ends up crowding everything into one container's shell script or handing the ordering to an orchestrator outside the cluster, and both methods lose the point of failure. If you build the two manifests side by side in this lab, you come away with a hands-on sense of where that boundary lies.

Steps

  1. Create the /root/capa-wf/ directory and, in workflow.yaml, write apiVersion argoproj.io/v1alpha1, kind Workflow, metadata.name capa-build, and spec.entrypoint main. spec.templates must contain a template whose name is main.
  2. In the main template, create dag.tasks and put in three tasks, checkout, build, and test. Make build depend on checkout and test depend on build, and do not attach dependencies to checkout.
  3. Put a template named build in spec.templates, and declare a parameter named revision in inputs.parameters and an artifact named binary with path /out/app in outputs.artifacts.
  4. In /root/capa-wf/cronworkflow.yaml, write kind CronWorkflow, metadata.name capa-nightly, spec.schedule 0 3 * * *, spec.concurrencyPolicy Forbid, and spec.workflowSpec.entrypoint main.
  5. In /root/capa-wf/workflowtemplate.yaml, write kind WorkflowTemplate and metadata.name capa-common, and put a template named notify in it. Then add a notify task to the main dag of workflow.yaml, with templateRef.name as capa-common, templateRef.template as notify, and dependencies as test.
  6. Create the namespace capa-wf in the cluster and actually create the Job capa-build-job in it. spec.backoffLimit is 2, the Pod's restartPolicy is Never, and the container image is busybox:1.36.
  7. In the same namespace, actually create the CronJob capa-nightly-job. spec.schedule is 0 3 * * *, spec.concurrencyPolicy is Forbid, and spec.successfulJobsHistoryLimit is 1.
  8. In the same namespace, actually create the ConfigMap capa-wf-summary. Put the number of main dag tasks in workflow.yaml in the key dag-tasks, the schedule value of cronworkflow.yaml in the key cron-schedule, and capa-build-job in the key job-name.

Notes

The Workflow skeleton and entrypoint

Create the /root/capa-wf/ directory and, in workflow.yaml, write apiVersion argoproj.io/v1alpha1, kind Workflow, metadata.name capa-build, and spec.entrypoint main. spec.templates must contain a template whose name is main.

The entrypoint points to the name of a template in the templates array. If the name it points to does not actually exist, the workflow cannot even start.

Draw the dependency graph

In the main template, create dag.tasks and put in three tasks, checkout, build, and test. Make build depend on checkout and test depend on build, and do not attach dependencies to checkout.

Each item in dag.tasks has a name, a template (or templateRef), and dependencies. Do not attach dependencies to the starting node of the graph.

Parameters and artifacts

Put a template named build in spec.templates, and declare a parameter named revision in inputs.parameters and an artifact named binary with path /out/app in outputs.artifacts.

Parameters are strings and artifacts are files. An output artifact needs both a name and a path inside the container.

Periodic execution with a CronWorkflow

In /root/capa-wf/cronworkflow.yaml, write kind CronWorkflow, metadata.name capa-nightly, spec.schedule 0 3 * * *, spec.concurrencyPolicy Forbid, and spec.workflowSpec.entrypoint main.

A CronWorkflow holds a Workflow's whole spec under a single field. Finding the name of that field is half of this step.

Reusing a WorkflowTemplate

In /root/capa-wf/workflowtemplate.yaml, write kind WorkflowTemplate and metadata.name capa-common, and put a template named notify in it. Then add a notify task to the main dag of workflow.yaml, with templateRef.name as capa-common, templateRef.template as notify, and dependencies as test.

templateRef needs two values: which object (name) and which template (template) to use. This task is also part of the graph, so you must give it a prerequisite.

Actually deploy a Job that does the same work

Create the namespace capa-wf in the cluster and actually create the Job capa-build-job in it. spec.backoffLimit is 2, the Pod's restartPolicy is Never, and the container image is busybox:1.36.

A Job Pod cannot use Always as its restartPolicy. The retry count is set not on the Pod but by a single field of the Job spec.

Express the same schedule with a CronJob

In the same namespace, actually create the CronJob capa-nightly-job. spec.schedule is 0 3 * * *, spec.concurrencyPolicy is Forbid, and spec.successfulJobsHistoryLimit is 1.

A CronJob also has a concurrency policy. Also check that there are separate fields for the number of records to keep, one for success and one for failure.

A summary that links the files and the cluster

In the same namespace, actually create the ConfigMap capa-wf-summary. Put the number of main dag tasks in workflow.yaml in the key dag-tasks, the schedule value of cronworkflow.yaml in the key cron-schedule, and capa-build-job in the key job-name.

You have to extract values from the files made in the earlier steps. You can put the result extracted with yq directly in as the ConfigMap values, or you can count by hand.