TT Lab
Get started
Learn Learning paths Courses

Infrastructure as Code

From Imperative to Declarative

Continue in TT Lab

In one line

The declarative approach defines the desired end state and lets the tool calculate the difference from the current state and carry out the changes, while the imperative approach describes the steps to run in order. In one sentence: you define the what of the infrastructure, and the tool handles the how.

Why this was needed

The problems with manual management boil down to four. You cannot recreate the same environment (not reproducible), no record remains of who changed what and when (changes cannot be traced), hands cannot keep up as the number of machines grows, and environments drift slightly apart from one another. The third is especially decisive. You can manage up to 10 servers by hand, but 100 or more is practically impossible.

Servers touched bit by bit by hand like that all become unique, like snowflakes. The metaphor of the Snowflake Server comes from here. The problem comes when such a server dies. Nobody can recreate that server exactly. That is because the configuration lived not in code but only on that machine's disk.

Imperative scripts do not fully solve this problem either. An order such as "install the package, copy the configuration file, and restart the service" is right the first time you run it, but what happens on the second run you can know only by reading the whole script. The declarative approach removes that question altogether. You write only the target state, and the tool calculates the difference from the current state.

How it works

A declarative tool always splits into three pieces. The desired state (code), the current state (read from the actual infrastructure), and the difference between the two (the plan). The difference is shown with symbols so people can read it, and the standard symbols are these. + create is something to create, - destroy is something to delete, ~ update is something to fix in place by changing only its attributes, -/+ replace is deletion followed by recreation, and <= read is a lookup that only reads. The symbol to watch most carefully in review is -/+. It reveals cases where you thought you were changing one attribute but the resource is recreated entirely.

In automation, you receive the result of the plan as an exit code. With plan -detailed-exitcode, 0 means no changes, 1 means an error, and 2 means there are changes to apply. You need these three values to build a pipeline such as "if there are changes, go to the approval stage; if not, just pass." If you save the plan result to a file and apply with that file, then even if the code changes between the time of the plan and the time of the apply, only the changes from the plan time are applied. The guarantee that what was reviewed and what was applied are the same comes from here.

When talking about drift, a three-axis model is handy. The code is Desired, the state file is Last Known, and the actual infrastructure is Actual. It is normal only when all three are the same, and if even two of them diverge, that is drift. GitOps tools keep running this comparison to self-heal. If someone changes replicas directly with kubectl, the controller detects it and reverts to the value defined in Git.

What you see in the field

The mistake newcomers make most often is to fix something hurriedly in the console and not reflect it in the code. At that moment the code becomes a lie, and the next person's apply silently reverts that change. Conversely, teams that run well attach the plan output to the PR for review. The key point is that what people review is not the code but the list of changes the code will produce.

The state file is the truly heavy asset

The place where declarative tools most often have incidents is not the code but the state file. It is the thing that plays "Last Known" in the three-axis model above. If this file is missing or broken, the tool does not recognize the existing resources as its own, and if you apply as it is, it tries to create what already exists again or to delete what belongs to someone else.

So a few things are basic.

Keep it in a shared store and lock it. If the state file is on each person's laptop, two people applying at the same time overwrite each other's records. Putting it in a remote store and locking it during apply is standard, and a setup with no lock is safe only by chance when there are few people.

Secrets end up inside it. Values you supplied when creating resources, such as a database password, often remain in the state file in plain text. So the state file is not committed to the code repository, is kept in an encrypted store, and has its access permissions managed differently from the code.

Do not edit it by hand. Move and delete entries only with the commands the tool provides. If you open it as text and edit it, the format is right but the internal references go out of step, and after a few applies an unexplained recreation occurs.

And splitting scope becomes more important as things grow. If you manage the whole infrastructure as a single state, a plan takes minutes and one small change locks everything. The principle is to split things with different lifetimes and owners. If you put things that almost never change like the network, things that change occasionally like the cluster, and things that change daily like the application in one lump, what changes daily exposes even what almost never changes to risk every time.

What you will do in the next lab

This Pod has no terraform binary. So you declare the desired state in YAML, read the actual list of containers as JSON in the same format, compare the two, and build your own plan that prints + create / - destroy / ~ update. Then you converge it with apply, delete a container from outside to create drift, and check that the plan catches it. Once you know how those symbols the tool prints on screen are produced, other people's plan output looks different to you too.