TT Lab
Get started
Learn Learning paths Courses

GitOps and Argo CD

Where Do the Buttons Come From

Continue in TT Lab

In one sentence

Argo CD's resource actions are built from two pieces of Lua written in argocd-cm — discovery.lua, which decides what to show, and action.lua, which decides what to change — and they can be tested without a server with argocd admin settings resource-overrides.

Why buttons were needed

The promise of GitOps is "the only way to change the cluster is the repository". But in the field, this promise is not kept perfectly. There are times when the configuration is baked into a cache and the Pods have to be brought up again once, times when you have to pause a deployment for a moment, and times during an outage when you have to raise something by one notch. These things have something in common — it is not that a value to be written in the repository changes. If you demand a commit when there is nothing to commit, people end up going to kubectl.

Resource actions fill this gap. You turn the few things that remain into predefined transformations and offer them as buttons on the screen, decide with RBAC who can press those buttons, and leave a record of who pressed. The key point is that only defined actions become possible, instead of arbitrary kubectl.

How it works

Two pieces go into one key of argocd-cm.

resource.customizations.actions.example.com_Widget: |
  discovery.lua: |
    actions = {}
    actions["pause"] = {}
    actions["resume"] = {["disabled"] = true}
    if obj.spec ~= nil and obj.spec.paused == true then
      actions["pause"] = {["disabled"] = true}
      actions["resume"] = {}
    end
    return actions
  definitions:
  - name: pause
    action.lua: |
      obj.spec.paused = true
      return obj

discovery.lua decides which buttons to show for this resource. It returns a table keyed by action names, and if you put disabled as true in a value, that button is dimmed. There is one important design decision here — the judgment of turning buttons on and off by looking at the state is in this code, not the screen. A button that stops something already stopped simply never becomes pressable in the first place.

In definitions, the action.lua receives the resource and returns the changed resource. It does not make a new object but edits the obj that came in and returns it, so it is easy to touch unintended fields along the way. That is why it matters that run-action does not show the result whole but shows only the changed fields as a diff — the places the action touched are visible at a glance.

There are two things in Lua that often trip you up. First, array indices start at 1. The first container is containers[1]. Second, reading a field that does not exist gives nil, and doing arithmetic on nil kills the script. An action that reads a value and computes needs a check.

Finally, you should know one property of this CLI. list-actions and run-action read only argocd-cm. Built-in actions that come inside Argo CD, such as a Deployment's restart, do not show up in this command, and if there is no configuration, it answers "no actions are configured". It is nothing to be alarmed about for being different from the screen; read it as meaning that this command tests only the rules I wrote.

What you see in the field

The first misunderstanding a team that adopts actions runs into is "if I press the button, it gets fixed". An action only changes objects in the cluster and does not change the repository. If automatic synchronization and self-healing are on, it goes back at the next reconcile. So actions are used only for things that are fine to go back — bringing Pods up again is the typical example. Changes that must not go back must still go through a commit.

The second is the quiet failure. A field name changed when the operator was upgraded, but action.lua stayed the same. The button is still visible, and when pressed, nothing happens or a value gets written in the wrong place. The side where no error occurs is the more dangerous. That is why actions too need the habit of bundling samples and expectations into a table — the last step of this lab is exactly that.

The limits of this lab environment

The lab Pod has neither an Argo CD controller nor a screen. You cannot press a button for real here, or block an action's permission with RBAC and watch it get refused. Instead, the code that actually runs discovery.lua and action.lua is inside the CLI, so you can confirm exactly the same result for which buttons are visible and what changes. And the changes the action produces are put directly on the kwok cluster so that you see the result with your own eyes.

What you will do in the next lab

You first check with an empty ConfigMap what this command reads. Then you make pause for a widget, make it dim alternately with resume depending on the state, and add scale-up, which reads a value and computes. You attach an action to the built-in kind Deployment too to pin the image, and apply that result to the kwok cluster to confirm it. Finally, you build a script that bundles what the actions do into a table and checks it all at once.