Who Turns the Green Light On
In one sentence
Argo CD's health is produced by a different verdict rule for each kind, a CRD with no built-in rule must have one written by hand in Lua in argocd-cm, and that rule can be tested on the spot, without a server, with argocd admin settings resource-overrides health.
Why health is separate from synchronization
The axis you meet first when learning GitOps is synchronization — is what is written in the repository the same as what is in the cluster. But there is one question this axis alone cannot answer. It is the same, but is it running well? Even if the image tag written in the repository does not exist, the Deployment object itself is created exactly like the repository. The synchronization axis is green and not a single Pod comes up. So Argo CD divides the axes in two. Synchronization is "are the declaration and the real thing the same", and health is "is the real thing doing its job". Only with the two axes separate can you express "Synced but Degraded", the state you meet most often in the field.
There are six health values — Healthy, Progressing, Degraded, Suspended, Missing, and Unknown. The health of the whole app is decided by collecting the values of the child resources, going with the worst. So even if the rule of a single child is wrong, the whole app looks the wrong color.
How it works
For built-in kinds such as Deployment, StatefulSet, Service, Ingress, Job, and PVC, the verdict code already exists inside Argo CD. For example, a Deployment looks at status's updatedReplicas, readyReplicas, and availableReplicas and at observedGeneration, and if it has not fully come up yet, it produces Progressing along with a message of "how many out of how many".
A CRD has no such code. So you lay a rule on argocd-cm. The key name is important.
data:
resource.customizations.health.example.com_Widget: |
hs = {}
if obj.status ~= nil and obj.status.phase == "Ready" then
hs.status = "Healthy"
hs.message = "widget is ready"
return hs
end
hs.status = "Progressing"
hs.message = "waiting for widget"
return hs
It is <그룹>_<종류> (group, then kind) and the separator is an underscore. Here, if you write the kind name differently from the manifest's kind, no error occurs — it simply behaves as if there were no rule. This is why this kind of mistake survives for a long time.
The Lua snippet receives the whole resource as a global variable called obj and returns a table with status and message filled in. The thing most often left out when writing a rule is the moment when status does not exist yet. A resource that was just created and has not been touched by a controller has no status at all. If you give Degraded here, newly created resources start with a red light every time, and the team that set up alerts soon turns those alerts off. The default branch must be Progressing.
Suspended is worth a separate mention. It is the slot that points to things a person has deliberately stopped — a pipeline under maintenance, a stopped CronJob. If you give Degraded for this, the pager rings for every planned maintenance. And the order of the branches changes the result. If the status.phase of a stopped resource is still Ready, a rule that wrote the Ready branch first ends up giving a green light.
What you see in the field
The most common incident is "a CRD that is always green". A database resource made by an operator has been failing to provision for hours, yet there is no indication on the screen. It is not that Argo CD lied, but that nobody wrote a rule to judge it, so that resource simply did not take part in the health calculation. Usually it is the users who raise the alarm first.
The second most common is a rule quietly going stale. When upgrading the operator, status.phase changed to status.conditions, but the Lua in argocd-cm stayed the same. The rule now always takes the default branch, and every resource freezes at Progressing. Nobody sees an error. That is why, once you write a rule, bundling sample resources and expected verdicts into a table and keeping it in the repository is as important as the rule itself. If you run this table once in CI, a red light appears right away on the day a field name changes.
The limits of this lab environment
The lab Pod has no Argo CD controller. So you cannot see "I applied the rule and the color on the screen changed", and you cannot confirm controller behaviors such as synchronization stopping because health is Degraded. Instead, the very code that actually performs the verdict is inside the CLI, so you can get the same answer from just one argocd-cm and one resource YAML.
What you will do in the next lab
You first look at the built-in verdict and check what sentence a CRD with no rule gives. Then you extend the Lua one branch at a time to build all four slots, Healthy, Degraded, Progressing, and Suspended, and see with your own eyes what happens when a key name is wrong by one character. Finally, you build a check script that bundles samples and expectations into a table and runs them all at once.