Three shards, yet only one pod ran
Goal
With Argo Workflows, actually run a pipeline in which the data produced by earlier steps decides the shape of later steps. You confirm from the node records how to fan out on script output, branch by condition, gather scattered results, absorb failures, wait for approval, and place a lock between workflows.
Why it matters
A data processing pipeline does not know how many pieces it will process until it runs. The step that splits the input must run first to determine the number of pieces, and then as many parallel tasks as that number must be created. You cannot express this by writing a list in YAML in advance, so Argo Workflows provides withParam, which uses the output of one node as the loop list of the next, when, which is evaluated per element, and automatic aggregation of fan-out output. In exchange, parallelism is the default, so for steps that use shared resources you must set limits and locks yourself, and you must also decide by design whether the failure of one piece counts as a failure of the whole. The exam asks what these template kinds and spec fields change in execution.
Steps
- Write a Workflow in
/root/capa-data/split.yaml, submit it to the argo namespace, and wait until it finishes. It has generateNamesplit-, the labelcapa-data: split, and an entrypoint DAGmainin which the script templatesplit(imagealpine:3.20, command[sh]) prints only a single line of the JSON list[{"id":"a","size":3},{"id":"b","size":12},{"id":"c","size":7}]to standard output, and then a taskpeekreceives that result as the parametershardsand prints it with busybox:1.36. - Write and submit a Workflow in
/root/capa-data/fanout.yaml(generateNamefanout-, labelcapa-data: fanout). In the entrypoint DAGmain, after the tasksplit(the script template from step 1), a taskprocessmust receive the result list ofsplitwith withParam and, for each element, run the container templateprocess(busybox:1.36) with the parametersidandsize. The workflow must be Succeeded. - In
/root/capa-data/serial.yaml, make the same fan-out Workflow as in step 2 with generateNameserial-and the labelcapa-data: serial, but have theprocesscontainer dosleep 3and putparallelism: 1on the DAG templatemain, and submit it. The run intervals of the threeprocessPods must not overlap. - Write and submit a Workflow in
/root/capa-data/branch.yaml(generateNamebranch-, labelcapa-data: branch). Aftersplit, two tasksbigandsmallboth receive the same list with withParam, andbiguseswhento run the container templatework(parametersidandlane) only for elements whose size is greater than 10, andsmallonly for elements whose size is 10 or less. The nodes of elements whose condition is false must remain as Skipped. - Write and submit a Workflow in
/root/capa-data/agg.yaml(generateNameagg-, labelcapa-data: agg). The fan-out taskcountcounts, for each element, the number of lines ofseq 1 <size>into a file and outputs it as the output parameterlines(valueFrom.path), and the tasktotalreceives{{tasks.count.outputs.parameters.lines}}as the parametervalues, computes the sum with a script, writes it to a file, and outputs it as the output parametersum. The output parameter sum of thetotalnode must be the sum of the three sizes. - Write and submit a Workflow in
/root/capa-data/tolerate.yaml(generateNametolerate-, labelcapa-data: tolerate). The fan-out taskcheckruns a container that fails with exit code 3 if size exceeds 10, but absorbs the failure withcontinueOn, and then a taskreportruns. The node of piece b must be Failed, and the workflow and report must be Succeeded. - Write a Workflow in
/root/capa-data/approve.yamlwith steps in the ordersplit→ the suspend templateapprove→ the container templatepublish(generateNameapprove-, labelcapa-data: approve) and submit it without--wait. Confirm that the workflow has stopped at approve, wait at least 10 seconds, then resume it withargo resumeso that it ends as Succeeded. - Create a ConfigMap
capa-data-locks(keywarehousewith value"1") in the argo namespace, and in/root/capa-data/locked.yamlwrite a Workflow that uses that key as a semaphore via spec.synchronization (generateNamelocked-, labelcapa-data: locked, the container runssleep 8). Submit two workflows in a row from the same file so that both end as Succeeded. The run intervals of the two workflows must not overlap.
Notes
- Inside the VM there are k3s, Argo Workflows v4.1.3 (in the argo namespace), and the argo CLI. Workflows run under the default account of the argo namespace.
busybox:1.36andalpine:3.20have been pulled in advance. - Submit and wait:
argo submit -n argo <파일> --wait, view nodes:argo get -n argo <이름>, list:argo list -n argo -l capa-data=<값>(the placeholders are the file, the workflow name, and the label value). - The grader looks at the most recent workflow for each step's label. If you fix it and submit again, the new workflow is graded.
- Common mistake: mixing
{{steps...}}and{{tasks...}}in withParam. Inside a DAG it is tasks, and inside steps it is steps. - Common mistake: comparing strings in a when expression without quotation marks. Number comparisons stay as they are, and string comparisons are wrapped in single quotes.
- Loops(withParam) · Conditionals · Suspending · Synchronization · Field Reference(continueOn)
A script's standard output becomes data
Write a Workflow in /root/capa-data/split.yaml, submit it to the argo namespace, and wait until it finishes. It has generateName split-, the label capa-data: split, and an entrypoint DAG main in which the script template split (image alpine:3.20, command [sh]) prints only a single line of the JSON list [{"id":"a","size":3},{"id":"b","size":12},{"id":"c","size":7}] to standard output, and then a task peek receives that result as the parameter shards and prints it with busybox:1.36.
A script template makes the source into a file and runs it with command, and puts the standard output out as the node's outputs.result. A later task reads it with {{tasks.<이름>.outputs.result}} (the placeholder is the task name). In this version (v4.1.3), the result of a script that nobody references was not left in the node record (measured). If you submit with argo submit --wait, it waits until the workflow finishes.
As many Pods are created as the length of the earlier list
Write and submit a Workflow in /root/capa-data/fanout.yaml (generateName fanout-, label capa-data: fanout). In the entrypoint DAG main, after the task split (the script template from step 1), a task process must receive the result list of split with withParam and, for each element, run the container template process (busybox:1.36) with the parameters id and size. The workflow must be Succeeded.
You read a DAG task's result with {{tasks.<이름>.outputs.result}} (the placeholder is the task name). The fields of a withParam element are {{item.<필드>}} (the placeholder is the field name). withItems is a method of writing the list in YAML in advance, so the length of the list is not decided at run time.
Let the fan-out flow through one at a time
In /root/capa-data/serial.yaml, make the same fan-out Workflow as in step 2 with generateName serial- and the label capa-data: serial, but have the process container do sleep 3 and put parallelism: 1 on the DAG template main, and submit it. The run intervals of the three process Pods must not overlap.
parallelism can be put on the whole workflow spec or on a single template. The grader looks at overlap using the nodes' startedAt and finishedAt.
Send to different branches by size
Write and submit a Workflow in /root/capa-data/branch.yaml (generateName branch-, label capa-data: branch). After split, two tasks big and small both receive the same list with withParam, and big uses when to run the container template work (parameters id and lane) only for elements whose size is greater than 10, and small only for elements whose size is 10 or less. The nodes of elements whose condition is false must remain as Skipped.
when evaluates as an expression the string after parameter substitution is finished. If you put withParam and when together on one task, they are evaluated separately for each element.
Gather the scattered results and add them up
Write and submit a Workflow in /root/capa-data/agg.yaml (generateName agg-, label capa-data: agg). The fan-out task count counts, for each element, the number of lines of seq 1 <size> into a file and outputs it as the output parameter lines (valueFrom.path), and the task total receives {{tasks.count.outputs.parameters.lines}} as the parameter values, computes the sum with a script, writes it to a file, and outputs it as the output parameter sum. The output parameter sum of the total node must be the sum of the three sizes.
If a later task reads the output parameter of a fan-out task, the per-element values are gathered into a JSON list string. Just pick out the numbers from that string and add them. The grader looks at the list that total received as well as the sum.
Absorb the failure of one piece and get through to the report
Write and submit a Workflow in /root/capa-data/tolerate.yaml (generateName tolerate-, label capa-data: tolerate). The fan-out task check runs a container that fails with exit code 3 if size exceeds 10, but absorbs the failure with continueOn, and then a task report runs. The node of piece b must be Failed, and the workflow and report must be Succeeded.
retryStrategy retries the same work, and continueOn acknowledges the failure and moves on. Compare what happens to the dependent tasks of the DAG when one element of a fan-out fails.
Stop and stand still until a person approves
Write a Workflow in /root/capa-data/approve.yaml with steps in the order split → the suspend template approve → the container template publish (generateName approve-, label capa-data: approve) and submit it without --wait. Confirm that the workflow has stopped at approve, wait at least 10 seconds, then resume it with argo resume so that it ends as Succeeded.
If you give a suspend template no duration, it waits until a person resumes it. While it is stopped, check the workflow's phase and the approve node's phase. The grader looks at how long the approve node stayed.
Keep two workflows from writing to the same warehouse at the same time
Create a ConfigMap capa-data-locks (key warehouse with value "1") in the argo namespace, and in /root/capa-data/locked.yaml write a Workflow that uses that key as a semaphore via spec.synchronization (generateName locked-, label capa-data: locked, the container runs sleep 8). Submit two workflows in a row from the same file so that both end as Succeeded. The run intervals of the two workflows must not overlap.
parallelism is a limit within one workflow, and a semaphore is a lock shared between workflows. While the second workflow waits, check status.synchronization and message. This version (v4.1.3) rejects the old singular field semaphore as an unknown field and accepts only the list fields (semaphores and mutexes) (measured).