TT Lab
Get started
Learn Learning paths Courses

CGOA — GitOps Certified Associate

Argo CD Is Green, but Checkout Returns 500

Continue in TT Lab

Goal

You separate Git synchronization from business success, check even the moment the configuration is consumed, and recover with Git.

Why it matters

Even when the desired configuration has been deployed, that configuration itself may be wrong, or the running process may not have read it. Using a real Argo CD, k3s, and nginx, you investigate a synthetic failure in which readiness is 200 and the business path is 500. You do not call or modify external payments, user data, or the production LabHub. Use only the cgoa-business-health Namespace on your personal VM. Do not change global controller, security, or network settings. It is a 55-minute lab. If needed, extend the time before it expires and keep the files you need separately. When the session ends, the VM and files are reclaimed.

Prepared environment and helper

The Git working repository is /srv/cgoa-business-health, and the bare repository is /srv/bare/cgoa-business-health.git. Argo CD reads the app directory of main from the gitd service inside the VM. It does not use an external Gitea or production Git. The lab server runs as non-root, with all capabilities dropped, a read-only root filesystem, RuntimeDefault, and no service account token mounted. The nginx image is pulled while the VM is being prepared, so you do not wait for a download in the first learning step. Student files go in /root/cgoa-health. The key=value notation in the tasks is explanatory; write the files as JSON, like the format examples. python3 /opt/fixtures/cgoa_health_lab.py observe reads the current Git, API, mounted file, and response. complete N verifies the answer you wrote for step N and saves the actual observation. It performs a bounded Git change only in steps 4 and 6. Check the changes yourself with git show --stat and git show. If you modified the repository directly, the helper stops without overwriting it. solve N is the same as the solution view and fills in a current answer that is missing. It does not modify an existing wrong answer or a partial answer. prepare N prepares only the earlier steps that are missing. It does not create the answer for the current step N. grade N only reads files, Git, and Kubernetes.

Steps

  1. On your personal VM, in /srv/cgoa-business-health, check git rev-parse HEAD. In revision.json, write the actual SHA string in revision and run complete 1. In observation-1.json, the local Git, remote main, and Argo CD revisions match, and the resource identities are preserved.
  2. With observe, read the responses of /healthz and the business /. In http.json, write the numbers readiness=200 and business=500, and run complete 2. observation-2.json keeps three pairs of actual HTTP codes and body samples. Do not interpret Healthy as business success.
  3. Compare the nginx.conf in git show HEAD:app/resources.json with the actual response. In diagnosis.json, write the strings cause=application-config, probe_scope=process, and fix_source=git, and run complete 3. Do not write an external API or DNS as if it were the actual cause.
  4. In config.json, write the number status=200 and the string body=checkout ready, and run complete 4. The helper creates a Git commit that modifies only the ConfigMap and pushes it to main. In observation-4.json, investigate whether the new revision and new ConfigMap are reflected but the previous Pod and the 500 response persist. Also check with git show that the Deployment has not changed.
  5. Compare the configmap.data.nginx.conf in observation-4.json with mounted_config and the previous Pod UID. In consumption.json, write the string mode=subPath and the booleans pod_replaced=false and restart_required=true, and run complete 5. Answer with the action required under this lab's way of replacing a new Pod. Do not expect the subPath to update just by waiting.
  6. Encode the bytes of configmap.data.nginx.conf in observation-4.json as UTF-8 and compute the SHA-256. In the checksum of rollout.json, write the computed 64-character string, and run complete 6. The second Git commit changes only the checksum/config in spec.template.metadata.annotations. The new Pod, the mounted file, and the business 200 response are preserved in observation-6.json.
  7. Investigate git_revision and pod.metadata.uid in observation-6.json. In verification.json, write revision and pod_uid as the corresponding strings, business=200 as a number, and deployment_replaced=false as a boolean, and run complete 7. Observe three actual response pairs again, and check that the Application and Deployment UIDs were kept.
  8. Based on the earlier observations and the Git lineage, write in decision.json the booleans recovered=true and availability_guaranteed=false and the strings config_strategy=versioned-rollout and next_check=external-path. Preserve the final observation with complete 8. Distinguish the success of an internal sample from a guarantee of external DNS, TLS, authentication, and long-term availability.

Notes

Look together not only at the Pod name but also at the UID, the ConfigMap data and the mounted file, and the HTTP code and body. The parent relationship between the Git configuration commit and the template commit remains in the observation's git_parents. Do not edit a successful observation. checksum/config is a convention for changing the template, not a built-in Kubernetes feature that automatically watches configuration. The three internal HTTP responses are a learning sample. They do not verify the external user path or long-term availability. Partial input is preserved. If there is a wrong answer, compare it with the task, fix it yourself, and then run complete again. If it is interrupted during a Git operation and a pending record is left, it does not automatically create a past success. Keep the material and start over in a new lab. The hashes detect accidental overwriting of records, but they are not a security guarantee against wholesale forgery by root on the same VM. Stay within the budgets of 60 seconds for grading and 90 seconds for preparing earlier steps. Waiting for actual convergence happens in complete, not in grading. Official ConfigMap · Argo CD health.

Linking Git to the actual synced commit

On your personal VM, in /srv/cgoa-business-health, check git rev-parse HEAD. In revision.json, write the actual SHA string in revision and run complete 1. In observation-1.json, the local Git, remote main, and Argo CD revisions match, and the resource identities are preserved.

Compare the commit SHA, which does not move, instead of the name main.

Observing the green light and the business 500 at the same time

With observe, read the responses of /healthz and the business /. In http.json, write the numbers readiness=200 and business=500, and run complete 2. observation-2.json keeps three pairs of actual HTTP codes and body samples. Do not interpret Healthy as business success.

That an HTTP response came back and that the business succeeded are different facts.

Telling the cause apart from the source to fix

Compare the nginx.conf in git show HEAD:app/resources.json with the actual response. In diagnosis.json, write the strings cause=application-config, probe_scope=process, and fix_source=git, and run complete 3. Do not write an external API or DNS as if it were the actual cause.

Find the return value in the nginx configuration that this server actually reads.

Fixing only the configuration through Git

In config.json, write the number status=200 and the string body=checkout ready, and run complete 4. The helper creates a Git commit that modifies only the ConfigMap and pushes it to main. In observation-4.json, investigate whether the new revision and new ConfigMap are reflected but the previous Pod and the 500 response persist. Also check with git show that the Deployment has not changed.

Separate the configuration change from the rollout with git show and the Pod UID.

Proving the old file left by subPath

Compare the configmap.data.nginx.conf in observation-4.json with mounted_config and the previous Pod UID. In consumption.json, write the string mode=subPath and the booleans pod_replaced=false and restart_required=true, and run complete 5. Answer with the action required under this lab's way of replacing a new Pod. Do not expect the subPath to update just by waiting.

Compare whether the configuration stored in the API and the file inside the container have the same value.

Updating the Pod template with a configuration hash

Encode the bytes of configmap.data.nginx.conf in observation-4.json as UTF-8 and compute the SHA-256. In the checksum of rollout.json, write the computed 64-character string, and run complete 6. The second Git commit changes only the checksum/config in spec.template.metadata.annotations. The new Pod, the mounted file, and the business 200 response are preserved in observation-6.json.

Compute the hash over the actual nginx.conf string, including the newlines.

Verifying the recovery with the new process's response

Investigate git_revision and pod.metadata.uid in observation-6.json. In verification.json, write revision and pod_uid as the corresponding strings, business=200 as a number, and deployment_replaced=false as a boolean, and run complete 7. Observe three actual response pairs again, and check that the Application and Deployment UIDs were kept.

Compare the new Pod UID and the preserved Deployment UID together.

Reporting the range you checked and the remaining verification

Based on the earlier observations and the Git lineage, write in decision.json the booleans recovered=true and availability_guaranteed=false and the strings config_strategy=versioned-rollout and next_check=external-path. Preserve the final observation with complete 8. Distinguish the success of an internal sample from a guarantee of external DNS, TLS, authentication, and long-term availability.

Three ClusterIP samples cannot guarantee the external path or a month of availability.