When Plans Got Slow, People Started Skipping Them
Goal
You build a 300-item state and measure plan time yourself, confirm with saved plans how each of the four knobs, refresh, parallelism, target narrowing, and state splitting, changes time and plan content, and then leave the measurement records as a table and a conclusion.
Why it matters
A slow plan is not an inconvenience but a safety problem. If you must wait 3 minutes, people skip the plan for small changes, and at that moment the discipline of 'look at what will change before applying' disappears. But every knob for handling a slow plan costs you something. If you turn off refresh, you cannot see changes made outside; if you narrow the target, that plan does not represent the whole configuration; and if you split the state, you must connect the layers through outputs. So order matters — first measure, confirm what eats the time, then pick what you can afford to lose and pull the knob. The absolute times in this lab differ from a real cloud (this Pod's provider makes no remote calls). What you learn is not the numbers but how to measure and how each knob changes the plan's content.
Steps
- In
/root/tfa-perf/main.tf, declare withcountas manyrandom_pet.nasvar.sizeand alocal_file.nthat writes each one's name to/root/tfa-perf/out/<인덱스>.txt(where the placeholder is the index). Putsize = 150in/root/tfa-perf/terraform.tfvars, then init and apply. The state will hold 300 instances. - Create
/root/tfa-perf/measure.sh. It takes the first argument as a label, passes the remaining arguments straight totofu plan, measures the elapsed time in milliseconds, saves the plan to/root/tfa-perf/plans/<이름표>.tfplanand the output to/root/tfa-perf/plans/<이름표>.log(where the placeholder is the label), and writes the three columns<이름표>,<밀리초>, and<실행한 명령>(the label, the milliseconds, and the command that was run) separated by tabs into/root/tfa-perf/times.tsv. If you measure again with the same label, the line must be replaced rather than added. Then measure once with the labelfulland no options. - Delete
/root/tfa-perf/out/7.txtoutside the tool to create drift, then measure with the labelno-refreshadding-refresh=false, and next measure with the labelrefreshwith no options. The number of changes in the two plan files must differ. Finally, apply to restore the deleted file. - Measure with the label
par1adding-parallelism=1. The number of change items in the saved plan must equal that offull. - Measure with the label
targetadding-target=random_pet.n[0]. Only a few items should remain in the saved plan, and/root/tfa-perf/plans/target.logmust contain the warning the tool issued. - Put the same configuration in
/root/tfa-perf/split/aand/root/tfa-perf/split/b, and init and apply each withsize = 75(together it is the same 300 as at the start). Then, in each directory, run/root/tfa-perf/measure.shonce each with the labelssplit-aandsplit-b. - After saving one plan with the label
stale, change the state in/root/tfa-perfwithtofu apply -replace=random_pet.n[0] -auto-approve. Then try to apply the savedplans/stale.tfplanand save its output to/root/tfa-perf/stale.txt. At the end, the plan must be clean. - Create
/root/tfa-perf/report.md. In a table that begins with the header| label | ms | command |and its separator line, put every line oftimes.tsvin label order, and after the table write one linefastest: <가장 빠른 이름표>(where the placeholder is the fastest label). Do not make up numbers; copy them as they are from times.tsv.
Notes
- The Pod has OpenTofu 1.9.0 and local and random provider mirrors, so it runs without the internet.
- You can open a saved plan (-out) with tofu show -json and count the change items. Unlike time, this number is the same even on a different machine.
- Common mistake: piling up measurement records only with append so that the same label becomes several lines. Later you cannot tell which line is the latest.
- Common mistake: measuring once and drawing a conclusion. The first run is slow because the cache is empty — measure twice with the same label and use the second.
- Command: plan · Command: apply · Resource Addressing · terraform_remote_state
Build a size worth measuring
In /root/tfa-perf/main.tf, declare with count as many random_pet.n as var.size and a local_file.n that writes each one's name to /root/tfa-perf/out/<인덱스>.txt (where the placeholder is the index). Put size = 150 in /root/tfa-perf/terraform.tfvars, then init and apply. The state will hold 300 instances.
The reason to use count here is to pair up and multiply the two resources by referencing each other through the index. On real infrastructure, these 300 would be checked against a remote API every time — that checking is most of the plan time.
Build the measuring tool first
Create /root/tfa-perf/measure.sh. It takes the first argument as a label, passes the remaining arguments straight to tofu plan, measures the elapsed time in milliseconds, saves the plan to /root/tfa-perf/plans/<이름표>.tfplan and the output to /root/tfa-perf/plans/<이름표>.log (where the placeholder is the label), and writes the three columns <이름표>, <밀리초>, and <실행한 명령> (the label, the milliseconds, and the command that was run) separated by tabs into /root/tfa-perf/times.tsv. If you measure again with the same label, the line must be replaced rather than added. Then measure once with the label full and no options.
You get milliseconds with the %s%3N format of date. If you remove the old line with the same label and write a new one, the record stays clean even when you measure several times — if measurement records only pile up by append, later you cannot tell which line is the latest. The reason to save the plan with -out is so that you can look again later at 'what was in it.'
What gets faster and what do you lose when you turn off refresh
Delete /root/tfa-perf/out/7.txt outside the tool to create drift, then measure with the label no-refresh adding -refresh=false, and next measure with the label refresh with no options. The number of changes in the two plan files must differ. Finally, apply to restore the deleted file.
Refresh is the job of checking one by one whether what is written in the state is actually still the same. If you turn it off, it skips that checking entirely, so it is faster, but you cannot see changes that arose outside. Open the saved plan with tofu show -json and count the number of changes.
If you lower parallelism, only the time changes and the result is the same
Measure with the label par1 adding -parallelism=1. The number of change items in the saved plan must equal that of full.
Parallelism is the value that decides how many tasks (mainly remote calls) the tool runs at the same time. It is unrelated to the content of the plan, so the result is the same and only the time differs. Does raising it always make it faster? No — if the target API imposes rate limits, retries increase and it actually gets slower.
If you narrow the target, only that much of the plan remains
Measure with the label target adding -target=random_pet.n[0]. Only a few items should remain in the saved plan, and /root/tfa-perf/plans/target.log must contain the warning the tool issued.
If you narrow the target, the tool puts only that resource and what it depends on into the plan. So this plan does not represent the whole configuration, and the tool notifies you of that fact with a warning. The official documentation says to use this option only in exceptional situations such as recovering from mistakes.
Split the state in two and measure the same number again
Put the same configuration in /root/tfa-perf/split/a and /root/tfa-perf/split/b, and init and apply each with size = 75 (together it is the same 300 as at the start). Then, in each directory, run /root/tfa-perf/measure.sh once each with the labels split-a and split-b.
measure.sh records based on where it is located, so wherever you call it from, the records gather in one place. Even if the two times after splitting add up to more than the original single one, the key point is that a person waits for only their own one state — splitting reduces not the total time but one person's waiting time.
A saved plan cannot be used once the state has changed
After saving one plan with the label stale, change the state in /root/tfa-perf with tofu apply -replace=random_pet.n[0] -auto-approve. Then try to apply the saved plans/stale.tfplan and save its output to /root/tfa-perf/stale.txt. At the end, the plan must be clean.
A saved plan is a promise that 'I will make this change from that state at that time.' If the state changes afterward, the premise of the promise breaks, so the tool rejects the apply. When running plan and apply separately in CI, this rule is exactly the safety device.
Organize the measurement records into a table and write the conclusion
Create /root/tfa-perf/report.md. In a table that begins with the header | label | ms | command | and its separator line, put every line of times.tsv in label order, and after the table write one line fastest: <가장 빠른 이름표> (where the placeholder is the fastest label). Do not make up numbers; copy them as they are from times.tsv.
The conclusion line must come from the numbers you measured — the grader also reads times.tsv and does the same calculation. Which knob cut the most may differ from machine to machine, and that is exactly the meaning of 'measure and then choose.'