Turn the Reconcile Loop by Hand, Once Around
Goal
You build yourself what happens while the reconcile loop goes around once. You go as far as applying the desired state together with ownership marks, stripping out the fields that fluctuate, comparing without the fields you decided to ignore, picking out deletion candidates, and refusing an empty result.
Why it matters
In the gitops-manifest lab, you made drift by hand and reverted it by hand. If you ask just one question here, the last piece of GitOps comes out. Who types that apply. If a person types it, it is a well-organized deployment script, not GitOps. For a controller to type it for you, the criteria for judgment must exist as code. They are what to treat as fluctuating values, which fields to leave owned by others, what to mark as owned by this app, and what to do when the result is empty. The worst case of selfHeal is that a manual change disappears, and the worst case of prune is that data disappears, so the levels of risk differ. When you write that judgment as a script, what lies behind the names of the switches becomes clear.
Environment
This Pod has no ArgoCD controller. You can create an Application object, but it does not become Synced by itself. So this lab goes the way of writing down a declaration and building yourself a tool that moves according to that declaration. Field ownership, server-side apply, and scheduling verdicts are handled by the real apiserver that kwok brings up, so that part is the real thing. The working directory is /root/gitops-sh, and under it you use k8s/, desired/, bin/, and out/.
Steps
- Declare the switches and the fields to ignore in
Application shop. - Apply the declared state server-side together with ownership marks.
- Strip out the fluctuating fields with
bin/normalize.sh. - Compare without the fields you decided to ignore, using
bin/drift.sh. - Create drift, revert only one side, and write it in
out/selfheal.txt. - Pick out deletion candidates with
bin/prune.sh. - Make
bin/sync.shrefuse an empty result. - Create a state that is Synced but not healthy and summarize it in
out/status.txt.
Notes
- In step 2, do not declare
spec.replicas. If you declare a field you decided to ignore, ownership comes over to this side and it collides every time. - In step 5, if you leave out
--force-conflicts, the image does not come back. It is not a failure but a difference in field ownership, so read the message. - The refusal in step 7 must stop without deleting anything. The grader checks that no object disappeared in between.
- Give execute permission to the three scripts.
Declare how to run the loop
In /root/gitops-sh/k8s/application.yaml, in the argocd namespace, write the Application shop and apply it. Under syncPolicy.automated, set prune: true, selfHeal: true, and allowEmpty: false; in syncOptions, put PruneLast=true and ServerSideApply=true; and with ignoreDifferences, make it ignore apps/Deployment's /spec/replicas. The destination namespace is sh-lab.
This Pod has no ArgoCD controller, so this object does not become Synced by itself. Instead, use it as the place where you write down how to run the loop. The scripts you build in the later steps only need to move as written here. allowEmpty defaults to false, but in this lab you state it explicitly to show the intent. ignoreDifferences is a device that keeps fields legitimately owned by another controller, such as an HPA, from being reverted.
Apply the desired state together with ownership marks
In /root/gitops-sh/desired/, create deployment.yaml (named orders) and configmap.yaml (named orders-config). Both have the namespace sh-lab, and the argocd.argoproj.io/tracking-id annotation must start with shop:. Do not write spec.replicas in the Deployment. Then apply the directory as a whole with --server-side --field-manager=argocd-controller.
The standard practice is not to declare at all a field you decided to ignore. If you declare it, this side takes the ownership of that field, and it collides every time with a value an HPA or a person changes later. tracking-id is the annotation with which ArgoCD marks what it owns, and it takes the form <앱>:<그룹>/<종류>:<네임스페이스>/<이름> (app, group/kind, namespace/name). Without the ownership mark, that object is not seen in the deletion-candidate calculation of step 6. After applying, look at who owns which field with kubectl get deploy orders -o yaml --show-managed-fields.
Build the preprocessing that strips out fluctuating fields
Create /root/gitops-sh/bin/normalize.sh <매니페스트> (the placeholder stands for the manifest). It prints to standard output the YAML from which, in metadata, resourceVersion, uid, generation, creationTimestamp, and managedFields, and the top-level status, are stripped out. What a person declared (kind, name, namespace, annotations, spec) must be left as it is.
If you compare as they are the values that Kubernetes fills in by itself and that keep changing, a difference is reported every time even if you change nothing. Then the reconcile loop never stops. Conversely, if you delete too much, you cannot see real differences either, so what to keep is as important as what to delete. You can read with python3's yaml module and print with yaml.safe_dump_all. The grader actually runs this script with a fixture.
Compare without the fields you decided to ignore
Create /root/gitops-sh/bin/drift.sh <선언파일> <실제파일> (the placeholders stand for the declared file and the actual file). After normalizing, compare without apps/Deployment's spec.replicas, and if they are the same, print SYNCED and end with 0, and if different, print OUTOFSYNC and end with 1.
The comparison goes in the direction of looking at whether what was declared is contained in the actual as it is. The actual side has lots of extra fields that Kubernetes filled in, and if you count even those as differences, nothing will pass. So it is convenient to follow the declared side's keys one by one and check that the actual side has the same value. For lists, match even the length and order. The grader actually runs this script with three pairs.
Create drift and revert only one side
With kubectl scale, raise orders's replicas to 5, and with kubectl set image, change the image tag to a different value. Then apply the declaration directory again with --server-side --field-manager=argocd-controller --force-conflicts, and write the result in five lines in /root/gitops-sh/out/selfheal.txt. They are DRIFT_FIELD, IGNORED, LIVE_REPLICAS, IMAGE_RESTORED, and CONFLICT_RESOLUTION.
Two things show up here at the same time. spec.replicas, which you did not declare, has a different owner and so stays as it is, and the image you declared comes back. If you leave out --force-conflicts, the image does not come back either. This is because kubectl set image took ownership of that field. The real ArgoCD also takes back ownership the same way when it synchronizes with ServerSideApply=true. Read LIVE_REPLICAS from the cluster and write it.
Pick deletion candidates by ownership mark
Create one ConfigMap in sh-lab without a declaration file and attach argocd.argoproj.io/tracking-id to it. Then create /root/gitops-sh/bin/prune.sh so that, for each object that is marked as owned by this app but is not in the declaration directory, it prints PRUNE=<종류>/<이름> (kind/name) on one line. It does not actually delete anything.
There are three criteria for the verdict. Read the declaration directory to build the list of what should exist, find in the cluster the objects this app owns, and pick from the second those that are not in the first. An object without an ownership mark is not a target no matter when it was made. You must not catch even things the cluster made by itself, such as kube-root-ca.crt. The grader does the same calculation itself and compares it with your output.
Stop when the render result is empty
Create /root/gitops-sh/bin/sync.sh. It reads the environment variable DESIRED_DIR (the default is the declaration directory), and if there is not a single manifest, it prints a message containing REFUSED and ends with a non-zero code. If there are, it applies with --server-side --field-manager=argocd-controller --force-conflicts, prints APPLIED=<개수> (the count), and ends with 0.
A single commit that wrongly changes source.path to an empty directory makes the list of what should exist 0 items, and all the objects that app was managing become deletion targets. It is enough for a reviewer to miss a one-character typo. That is why refusing the empty result itself is the most direct defense. When you refuse, you must stop without deleting anything. The grader runs it once with an empty directory and once with a normal directory, and also checks that no object disappeared in between.
Create a state that is Synced but not healthy
In /root/gitops-sh/desired/broken.yaml, declare a Deployment with a nodeSelector that cannot be scheduled, broken (including the ownership mark), and apply it with sync.sh. Then summarize it in five lines in /root/gitops-sh/out/status.txt. They are SYNC, HEALTH, REASON, PRUNE, and ALLOW_EMPTY.
The Sync status is whether it is the same as the repository and the Health status is whether it is running well, so they are different axes. That is why a combination in which it was deployed as the repository ordered and yet the Pod cannot come up appears normally. In practice, the case where the image tag was written wrong is exactly this state, and if you read the two axes mixed together, you look for the cause of the outage in the wrong place. If you write in nodeSelector a label that does not exist in the cluster, you can make the same state. For PRUNE and ALLOW_EMPTY, write the values you declared in step 1 as they are.