TT Lab
Get started
Learn Learning paths Courses

CAPA — Argo Project Associate

Where did the successful pipeline leave its files?

Continue in TT Lab

In one line

A Workflow's success, the cleanup of intermediate files, and the retention of the final result are different contracts. Only by checking each of the three contracts can you trust the green light of CI and move on to the next task.

Why this was needed

The nightly data job succeeded. But in the morning the team that reads the result cannot find the file. The person in charge increases the retry count and raises the container memory. Would that help? If the production job did not fail but the post-success cleanup policy also deleted the final result, it only succeeds faster and deletes the file more reliably. You have to tell which layer failed first for the direction of recovery to be right.

There is also the opposite incident. The result was produced correctly, but temporary files keep piling up. The Workflow list shows Succeeded, so no alert goes off. If the service account responsible for cleanup has lost its Kubernetes permissions, the exit code of the business container and the error of the GC can differ at the same time. In this module, you reproduce the two incidents separately using a real S3-compatible store and real containers.

How it works

The path is the boundary of execution

The producer writes to /work/temp.txt inside the Pod. The executor uploads that file to the object store under a key. The consumer downloads it again from the key in the store. A local file path and an S3 key are not the same concept. Two different Pods may safely use the same path locally, but if they upload to the same key in the same bucket, the later upload can overwrite the result written first.

generateName reduces collisions of Kubernetes object names. That alone does not make the S3 key different. If you put {{workflow.uid}} in the output key, this run's identifier carries through to the storage path. If you also put the producer UID in the file content and the consumer compares it with its own UID, you can do a stronger check than merely whether the file exists. It also detects the incident of reading someone else's valid file.

The two normal runs in this lab use runs/<실행 UID>/temp.txt and runs/<실행 UID>/final.json respectively (the placeholder is the run UID). The collision experiment has the two runs share the shared path of the same batch. So as not to delete the shared file of another experiment attempt, the experiment helper distinguishes the batch itself every time. The intended collision occurs only within the same attempt.

When to delete and what to keep

The Workflow-level spec.artifactGC.strategy is the default cleanup policy. OnWorkflowCompletion operates on completion, and OnWorkflowDeletion operates on Workflow deletion. The artifactGC.strategy: Never of an individual output artifact excludes that file from cleanup. You can express a policy that deletes intermediate products while keeping the final result.

Never here is not a backup. If the store disappears or an operator deletes the object, the file is gone. The lab store is a disposable emptyDir inside the VM, so it is not kept after the session ends. In production you must separately design persistent volumes, recoverable backups, access control, TLS, and retention periods.

Reading a Workflow's success and confirming file retention are also different. This lab compares the UID and hash of the bytes actually read, and when judging deletion, it checks an authenticated list query together with a 404 NoSuchKey. You must not conclude "the file was deleted" from a connection failure or a 403.

Kubernetes permissions and store permissions are separate

The GC Pod must first look at the WorkflowArtifactGCTask and record the processing result in its status. The Kubernetes permissions needed for this are list/watch on that resource and patch on the status subresource. After that, deleting an S3 object requires the permissions of the store account. If you fixed the Kubernetes Role and still get AccessDenied in S3, you are looking at a different layer. Do not mix the two permission systems.

You check by specifying the subresource explicitly, as in kubectl auth can-i patch workflowartifactgctasks --subresource=status. If you just throw in the argument workflowartifactgctasks/status, it is interpreted with a different meaning and you can see a wrong no. You must include in the evidence not only the output of the command but also what question was asked.

In this lab, the GC that deliberately has no permissions leaves ArtifactGCError and a forbidden log. The business status can still be Succeeded. Forcibly removing the error and the finalizer to make it green is not recovery. You preserve the original failure evidence, attach a minimal Role, and then verify with a new Workflow UID that cleanup and retention are normal. You do not assume that the original GC that already failed is automatically reprocessed.

What it looks like in the field

In model training, you delete large intermediate checkpoints but must keep the final model and the evaluation report. In data transformation, two overlapping reruns in the same time window must not delete each other's inputs. For a regulated report, it is better to manage "the job succeeded" and "the result can be read back and verified" as separate metrics. It also matters that the people responsible for responding to a failed job, a failed cleanup, and a wrong retention setting may not be the same.

In LabHub's own environment verification, there was a case where the service was confirmed Ready but a normal S3 upload returned 500. The cause was that the initial metadata took up the volume limit of the small lab first. If we had checked only that an anonymous request is a 403, we would have missed this failure. That is why you keep both a control group where a request that should be allowed succeeds and a counterexample where a request that should be rejected fails.

What you will do in the next lab

First you check the store's access boundary. You fix a shared key into a per-UID key and compare two runs, and then analyze separately the missing final retention and the GC permission failure. You verify a new run with minimal permissions, and observe the shared key collision and GC on deletion. At the end, you explain, by putting in the real UIDs, that the same Succeeded can have different storage results. The helper's completion message means the experiment run has finished and is not the same as passing the lab grading.

Official documentation