Trace missing files behind a green pipeline
Goal
Distinguish business success from the success of artifact retention and cleanup, and verify with collision-free keys and least privilege.
Why it matters
Even when the business succeeds, the final file can disappear or intermediate files can keep piling up. This lab uses the real Argo Workflows 4.1.3 and SeaweedFS 4.46 inside the VM. It checks, together, not the success mark on the screen but the run UID, the actual file bytes, and the GC errors. The default cluster and store are prepared automatically, and the student modifies policies or runs the comparison scenarios directly.
Steps
- Read the artifact-repositories and the Service seaweed in the argo namespace. Record namespace=argo, endpoint=seaweed:8333, writer_bucket=artifacts, and anonymous_allowed=false in /root/capa-artifacts/inventory.json. Grading also checks the actual file read by the permitted account and the rejection of an anonymous request and of access to another team's bucket.
- In the two produce outputs of the prepared /root/capa-artifacts/isolated.yaml, fix only s3.key. temporary is runs/{{workflow.uid}}/temp.txt and final is runs/{{workflow.uid}}/final.json. Keep the rest of the pipeline and the final Never, and run isolated with the experiment helper. Compare the UIDs of the two different Workflows and the final file contents.
- After the two isolated runs complete, check the store. Record business=Succeeded, temporary=deleted, final=retained, and final_policy=Never in /root/capa-artifacts/retention.json. Also check the experiment record to see that the production UIDs of the two final files are different and that B's input was still alive at the time of A's cleanup.
- The prepared missing-retain.yaml deliberately left out Never on the final artifact. After running missing-retain with the helper, record business=Succeeded, final_present=false, and fix=Never in /root/capa-artifacts/missing.json. Do not change the YAML of this comparison run to the normal policy. This step is for explaining the cause of the error and the field to fix.
- The prepared gc-denied.yaml uses the GC account gc-denied, which has no Role. Run gc-denied with the helper and write business=Succeeded, condition=ArtifactGCError, cause=RBAC, and force_finalizer=false in /root/capa-artifacts/gc-error.json. Read the actual forbidden log and the two remaining files, and do not remove the finalizer.
- Write and apply a RoleBinding capa-gc-repair in /root/capa-artifacts/gc-binding.yaml. The namespace is argo, roleRef is apiGroup=rbac.authorization.k8s.io, kind=Role, name=artifact-gc, and subjects is a single ServiceAccount gc-denied in argo. Do not widen the existing Role, and confirm a new run with gc-repaired.yaml. Preserve the record and finalizer of the original failed Workflow.
- Keep the shared/{{workflow.parameters.batch}} key in collision.yaml and run collision with the helper. Confirm that A's input is overwritten with B's UID and A fails, and then A's GC deletes B's input so that B fails too. The Failed/Error of both Workflows is the expected result of this comparison. Do not fake the failure by simply deleting the file.
- on-deletion.yaml uses OnWorkflowDeletion together with the final Never. Run on-deletion with the helper. The helper observes the files right after success, confirms the exact Workflow UID, deletes only that Workflow, and observes again. At business completion both files must be there, and after deletion only the final file must remain.
- In /root/capa-artifacts/report.json, write the actual UID of missing-retain for lost_result_uid and the actual UID of gc-denied for gc_failure_uid. Record same_business_status=true, same_storage_result=false, and status_is_retention_proof=false. Explain why the two Succeeded call for different actions, based on the actual file state and the error layer.
Notes
- The prepared YAML is in /root/capa-artifacts/. Apart from the two keys in step 2, do not remove the wrong settings needed for the comparison. Each step card has the exact changes and observations.
- Run: python3 /opt/fixtures/capa-lifecycle/runtime.py run isolated /root/capa-artifacts/isolated.yaml Other scenarios take the same form, and the name is the file name of that YAML minus .yaml.
- Status: python3 /opt/fixtures/capa-lifecycle/runtime.py status isolated
- The original before and after observations of the attempt that status shows are in
/var/lib/labhub/capa-lifecycle/attempts/<attempt>.json. before is after production, after is after consumption and cleanup, and while_a_finished is B's state when A finished. This is the helper's observation record, not the student's answer, so do not edit it directly. - To look up the actual run:
kubectl -n argo get workflow WORKFLOW_NAME -o yamlTo look up the actual Pods:kubectl -n argo get pods -l workflows.argoproj.io/workflow=WORKFLOW_NAMETo look up logs:kubectl -n argo logs POD_NAME --all-containers=trueReplace WORKFLOW_NAME and POD_NAME with the actual names found with the earlier command. - Store listing:
kubectl -n argo exec store-admin -- mc --config-dir /tmp/mc ls --recursive store/artifacts/Reading a file:kubectl -n argo exec store-admin -- mc --config-dir /tmp/mc cat store/artifacts/runs/<UID>/final.jsonstore-admin is an observation admin tool inside this VM. Business uploads and GC use a separate artifact-writer account, and step 1 checks that account's bucket restrictions. - A run can continue even after the observation time ends. Check the same job with wait isolated, and do not create a duplicate of a job that is still alive. The completion record of the same YAML is reused. Attach --new-attempt to run only when you intentionally want a fresh comparison of a finished experiment.
- The helper pauses after the upload, observes, and resumes consumption. The helper's completion of a run is not a pass of the correct answer. A run that ended with wrong YAML is rejected by grading.
- For kubectl, use KUBECONFIG=/etc/rancher/k3s/k3s.yaml. For the argo CLI, specifying --kubeconfig=/etc/rancher/k3s/k3s.yaml explicitly keeps it from depending on the home directory.
- The store's HTTP and emptyDir are disposable, for the lab. When the session ends, the files and the VM are both reclaimed. Never is also not a feature that keeps or backs up beyond the end of the session.
Distinguish a normal read from a read that should be rejected
Read the artifact-repositories and the Service seaweed in the argo namespace. Record namespace=argo, endpoint=seaweed:8333, writer_bucket=artifacts, and anonymous_allowed=false in /root/capa-artifacts/inventory.json. Grading also checks the actual file read by the permitted account and the rejection of an anonymous request and of access to another team's bucket.
The endpoint and the bucket in the repository reference are different fields. Also confirm that a normal request succeeds.
Separate the artifact keys of two runs
In the two produce outputs of the prepared /root/capa-artifacts/isolated.yaml, fix only s3.key. temporary is runs/{{workflow.uid}}/temp.txt and final is runs/{{workflow.uid}}/final.json. Keep the rest of the pipeline and the final Never, and run isolated with the experiment helper. Compare the UIDs of the two different Workflows and the final file contents.
Even if the Kubernetes names differ, if the store key is the same they collide. Put the UID in the S3 key, not in the local file name.
Delete intermediate files and keep the final result
After the two isolated runs complete, check the store. Record business=Succeeded, temporary=deleted, final=retained, and final_policy=Never in /root/capa-artifacts/retention.json. Also check the experiment record to see that the production UIDs of the two final files are different and that B's input was still alive at the time of A's cleanup.
Look at the default OnWorkflowCompletion policy together with the individual final's Never. You can check the observation record after finding the run name in the helper status.
A success analysis that missed the final retention
The prepared missing-retain.yaml deliberately left out Never on the final artifact. After running missing-retain with the helper, record business=Succeeded, final_present=false, and fix=Never in /root/capa-artifacts/missing.json. Do not change the YAML of this comparison run to the normal policy. This step is for explaining the cause of the error and the field to fix.
A wrong configuration may not produce a business failure. Compare the normal isolated result with this run's final file.
Separate business success from a GC permission failure
The prepared gc-denied.yaml uses the GC account gc-denied, which has no Role. Run gc-denied with the helper and write business=Succeeded, condition=ArtifactGCError, cause=RBAC, and force_finalizer=false in /root/capa-artifacts/gc-error.json. Read the actual forbidden log and the two remaining files, and do not remove the finalizer.
What must fail at this step is not the business container but the GC. Attaching permissions is done in the next step.
Verify a new run with minimal GC permissions
Write and apply a RoleBinding capa-gc-repair in /root/capa-artifacts/gc-binding.yaml. The namespace is argo, roleRef is apiGroup=rbac.authorization.k8s.io, kind=Role, name=artifact-gc, and subjects is a single ServiceAccount gc-denied in argo. Do not widen the existing Role, and confirm a new run with gc-repaired.yaml. Preserve the record and finalizer of the original failed Workflow.
GC needs only task list/watch and task status patch. Check the status permission with --subresource=status.
Two failures caused by a shared key
Keep the shared/{{workflow.parameters.batch}} key in collision.yaml and run collision with the helper. Confirm that A's input is overwritten with B's UID and A fails, and then A's GC deletes B's input so that B fails too. The Failed/Error of both Workflows is the expected result of this comparison. Do not fake the failure by simply deleting the file.
The helper gives the same batch only to the two runs of the same experiment attempt. Compare the different UIDs shown in status and the logs.
Compare completion time and deletion time
on-deletion.yaml uses OnWorkflowDeletion together with the final Never. Run on-deletion with the helper. The helper observes the files right after success, confirms the exact Workflow UID, deletes only that Workflow, and observes again. At business completion both files must be there, and after deletion only the final file must remain.
Deleting a Workflow and deleting a Pod are not the same event. Read the evidence that the deletion finished without forcibly removing the finalizer.
Report the different storage results hidden behind the same success
In /root/capa-artifacts/report.json, write the actual UID of missing-retain for lost_result_uid and the actual UID of gc-denied for gc_failure_uid. Record same_business_status=true, same_storage_result=false, and status_is_retention_proof=false. Explain why the two Succeeded call for different actions, based on the actual file state and the error layer.
Read the UIDs from the helper status, but do not confuse them with names. Losing the final result and failing to clean up temporary files are not the same failure.