TT Lab
Get started
Learn Learning paths Courses

Terraform in Practice

Four Levers for a Slow Plan, and What Each Costs

Continue in TT Lab

In one sentence

There are four knobs for reducing plan time — turning off refresh, parallelism, narrowing the target, and splitting the state. All four cost you something, so after measuring you must pick what you can afford to lose.

Why this is a problem — a slow plan is a safety problem

When a plan starts taking 3 minutes, people's behavior changes. Not wanting to wait 3 minutes after fixing a single line, they say "this is an obvious change" and go straight to apply. Once that habit settles in, the discipline of "look at what will change before applying" itself disappears. The incident comes after that.

So plan performance is not a matter of convenience but a matter of making it possible to keep the discipline. But you must not pull just any knob. You need to know exactly what each knob gives up.

How it works

A single plan does three big things. It reads the configuration and builds a graph (fast), checks one by one whether what is written in the state is actually still the same (usually most of the time is here), and compares the result with the configuration to build the list of changes (fast). The reason a large state is slow is almost always the middle.

1. Turning off refresh. It skips the middle stage entirely. It reduces the most, but you cannot see changes that arose outside. Even if someone edited it by hand in the console, the plan says "no changes." It is worth using only when you must revert in a hurry or when you have just applied directly and the state is certainly up to date.

2. Parallelism. It sets the number of tasks proceeding at the same time. The content of the plan does not change at all; only the time differs. Does raising it always make it faster? No — if the target API imposes rate limits, retries increase and it actually gets slower. There are also cases where you need to lower it (strict API limits, shared accounts).

3. Narrowing the target. It puts only the specified address and what it depends on into the plan. It is certainly faster, but that plan does not represent the whole configuration. The tool also issues a warning every time you plan. The OpenTofu documentation says to use this option "only in exceptional situations, such as recovering from mistakes or working around tool limitations." If you are using it in everyday work, that is a signal that it is not a performance problem but a design problem.

4. Splitting the state. This is the fundamental solution. If you put the layer that changes every day and the layer that rarely changes into different states, each person waits only for their own layer. The sum of total time may actually increase, but one person's waiting time decreases, and as a bonus lock contention and incident blast radius decrease too. The cost is that you must connect the layers through outputs, and that connection becomes a contract.

On top of this comes a saved plan (-out). You use it when running plan and apply separately in CI, and there is one rule. A saved plan presumes the state as it was at that moment. If the state changes in between, the apply is rejected. It may look inconvenient, but this is precisely the device that guarantees "the plan you reviewed is the one applied."

What it looks like in the field

The most common mistake is choosing without measuring. It starts with "it's slow, so let's turn off refresh," and half a year later you are in a state where drift has piled up and nobody knows. If you measure, the cause is usually more specific — a particular data source lists thousands of items every time, or one module takes up half of the state.

The second is measuring once and drawing a conclusion. The first run is slow because the cache is empty. You need the habit of measuring twice under the same conditions and using the second number.

The third is only piling up measurement records. If the same label piles up on several lines, later you cannot tell which line is the latest. Make the records idempotent too — make the same label overwrite.

And one fact that is often missed in practice. Time differs from machine to machine, but the content of the plan does not. If you open a saved plan as JSON and count the change items, the same number comes out on any machine. So claims such as "if you narrow the target, only that much of the plan remains" and "if you turn off refresh, you cannot see drift" should be proven not by time but by this number.

Fourth, a criterion that separates where a performance problem ends and a design problem begins makes conversation easier. If pulling one knob made it bearable, it is a performance problem. If you are in a state where work gets done only by narrowing the target in everyday work, it is already a design problem, and the answer is not an option but the work of splitting the state. If you do not make this distinction, the team ends up hardening a stopgap into a standard procedure.

Finally, do not forget the cost of measuring itself. If you measure a plan ten times, that costs that much time. So in practice you measure and record only one before-and-after pair. In the record, leave not only the number but the whole command you used — half a year later you will be asked "under what conditions did this number come from?" and without the command you cannot answer that, so you have no choice but to measure again.

What to do in the next lab

You build a 300-item state in /root/tfa-perf and first build a measuring script. Next you delete one file outside the tool to create drift and see how the number of changes in the plan differs with refresh off and on, measure parallelism and target narrowing, and measure again after splitting the state in two. Finally you confirm that a saved plan is rejected because of a state change, and leave the measured numbers as a table and a conclusion.