TT Lab
Get started
Learn Learning paths Courses

Terraform in Practice

When Plans Got Slow, People Started Skipping Them

Continue in TT Lab

Goal

You build a 300-item state and measure plan time yourself, confirm with saved plans how each of the four knobs, refresh, parallelism, target narrowing, and state splitting, changes time and plan content, and then leave the measurement records as a table and a conclusion.

Why it matters

A slow plan is not an inconvenience but a safety problem. If you must wait 3 minutes, people skip the plan for small changes, and at that moment the discipline of 'look at what will change before applying' disappears. But every knob for handling a slow plan costs you something. If you turn off refresh, you cannot see changes made outside; if you narrow the target, that plan does not represent the whole configuration; and if you split the state, you must connect the layers through outputs. So order matters — first measure, confirm what eats the time, then pick what you can afford to lose and pull the knob. The absolute times in this lab differ from a real cloud (this Pod's provider makes no remote calls). What you learn is not the numbers but how to measure and how each knob changes the plan's content.

Steps

  1. In /root/tfa-perf/main.tf, declare with count as many random_pet.n as var.size and a local_file.n that writes each one's name to /root/tfa-perf/out/<인덱스>.txt (where the placeholder is the index). Put size = 150 in /root/tfa-perf/terraform.tfvars, then init and apply. The state will hold 300 instances.
  2. Create /root/tfa-perf/measure.sh. It takes the first argument as a label, passes the remaining arguments straight to tofu plan, measures the elapsed time in milliseconds, saves the plan to /root/tfa-perf/plans/<이름표>.tfplan and the output to /root/tfa-perf/plans/<이름표>.log (where the placeholder is the label), and writes the three columns <이름표>, <밀리초>, and <실행한 명령> (the label, the milliseconds, and the command that was run) separated by tabs into /root/tfa-perf/times.tsv. If you measure again with the same label, the line must be replaced rather than added. Then measure once with the label full and no options.
  3. Delete /root/tfa-perf/out/7.txt outside the tool to create drift, then measure with the label no-refresh adding -refresh=false, and next measure with the label refresh with no options. The number of changes in the two plan files must differ. Finally, apply to restore the deleted file.
  4. Measure with the label par1 adding -parallelism=1. The number of change items in the saved plan must equal that of full.
  5. Measure with the label target adding -target=random_pet.n[0]. Only a few items should remain in the saved plan, and /root/tfa-perf/plans/target.log must contain the warning the tool issued.
  6. Put the same configuration in /root/tfa-perf/split/a and /root/tfa-perf/split/b, and init and apply each with size = 75 (together it is the same 300 as at the start). Then, in each directory, run /root/tfa-perf/measure.sh once each with the labels split-a and split-b.
  7. After saving one plan with the label stale, change the state in /root/tfa-perf with tofu apply -replace=random_pet.n[0] -auto-approve. Then try to apply the saved plans/stale.tfplan and save its output to /root/tfa-perf/stale.txt. At the end, the plan must be clean.
  8. Create /root/tfa-perf/report.md. In a table that begins with the header | label | ms | command | and its separator line, put every line of times.tsv in label order, and after the table write one line fastest: <가장 빠른 이름표> (where the placeholder is the fastest label). Do not make up numbers; copy them as they are from times.tsv.

Notes

Build a size worth measuring

In /root/tfa-perf/main.tf, declare with count as many random_pet.n as var.size and a local_file.n that writes each one's name to /root/tfa-perf/out/<인덱스>.txt (where the placeholder is the index). Put size = 150 in /root/tfa-perf/terraform.tfvars, then init and apply. The state will hold 300 instances.

The reason to use count here is to pair up and multiply the two resources by referencing each other through the index. On real infrastructure, these 300 would be checked against a remote API every time — that checking is most of the plan time.

Build the measuring tool first

Create /root/tfa-perf/measure.sh. It takes the first argument as a label, passes the remaining arguments straight to tofu plan, measures the elapsed time in milliseconds, saves the plan to /root/tfa-perf/plans/<이름표>.tfplan and the output to /root/tfa-perf/plans/<이름표>.log (where the placeholder is the label), and writes the three columns <이름표>, <밀리초>, and <실행한 명령> (the label, the milliseconds, and the command that was run) separated by tabs into /root/tfa-perf/times.tsv. If you measure again with the same label, the line must be replaced rather than added. Then measure once with the label full and no options.

You get milliseconds with the %s%3N format of date. If you remove the old line with the same label and write a new one, the record stays clean even when you measure several times — if measurement records only pile up by append, later you cannot tell which line is the latest. The reason to save the plan with -out is so that you can look again later at 'what was in it.'

What gets faster and what do you lose when you turn off refresh

Delete /root/tfa-perf/out/7.txt outside the tool to create drift, then measure with the label no-refresh adding -refresh=false, and next measure with the label refresh with no options. The number of changes in the two plan files must differ. Finally, apply to restore the deleted file.

Refresh is the job of checking one by one whether what is written in the state is actually still the same. If you turn it off, it skips that checking entirely, so it is faster, but you cannot see changes that arose outside. Open the saved plan with tofu show -json and count the number of changes.

If you lower parallelism, only the time changes and the result is the same

Measure with the label par1 adding -parallelism=1. The number of change items in the saved plan must equal that of full.

Parallelism is the value that decides how many tasks (mainly remote calls) the tool runs at the same time. It is unrelated to the content of the plan, so the result is the same and only the time differs. Does raising it always make it faster? No — if the target API imposes rate limits, retries increase and it actually gets slower.

If you narrow the target, only that much of the plan remains

Measure with the label target adding -target=random_pet.n[0]. Only a few items should remain in the saved plan, and /root/tfa-perf/plans/target.log must contain the warning the tool issued.

If you narrow the target, the tool puts only that resource and what it depends on into the plan. So this plan does not represent the whole configuration, and the tool notifies you of that fact with a warning. The official documentation says to use this option only in exceptional situations such as recovering from mistakes.

Split the state in two and measure the same number again

Put the same configuration in /root/tfa-perf/split/a and /root/tfa-perf/split/b, and init and apply each with size = 75 (together it is the same 300 as at the start). Then, in each directory, run /root/tfa-perf/measure.sh once each with the labels split-a and split-b.

measure.sh records based on where it is located, so wherever you call it from, the records gather in one place. Even if the two times after splitting add up to more than the original single one, the key point is that a person waits for only their own one state — splitting reduces not the total time but one person's waiting time.

A saved plan cannot be used once the state has changed

After saving one plan with the label stale, change the state in /root/tfa-perf with tofu apply -replace=random_pet.n[0] -auto-approve. Then try to apply the saved plans/stale.tfplan and save its output to /root/tfa-perf/stale.txt. At the end, the plan must be clean.

A saved plan is a promise that 'I will make this change from that state at that time.' If the state changes afterward, the premise of the promise breaks, so the tool rejects the apply. When running plan and apply separately in CI, this rule is exactly the safety device.

Organize the measurement records into a table and write the conclusion

Create /root/tfa-perf/report.md. In a table that begins with the header | label | ms | command | and its separator line, put every line of times.tsv in label order, and after the table write one line fastest: <가장 빠른 이름표> (where the placeholder is the fastest label). Do not make up numbers; copy them as they are from times.tsv.

The conclusion line must come from the numbers you measured — the grader also reads times.tsv and does the same calculation. Which knob cut the most may differ from machine to machine, and that is exactly the meaning of 'measure and then choose.'