Terraform/OpenTofu Fundamentals
How to Read a Plan — Symbols, Exit Codes, Drift
In one line
apply only carries out the decisions plan has already made. The only place to prevent an accident is the few minutes of reading the plan, and in the end those few minutes have to be taken over not by a person but by a script.
Why this was needed
The most expensive mistake in infrastructure code is "misreading a recreation as an in-place update." If you mistake the -/+ on the screen for ~, a production database is deleted and created again. Plan output is long, deployments usually happen late at night, and the human eye stops reading after seeing the same screen three times.
So mature teams do two things. First, they save the plan to a file and apply exactly the plan that was reviewed. If you do not save it, the plan reviewed and the plan actually applied can differ. Someone may have changed something else in between. Second, they extract the plan in a machine-readable format and judge it by rules. A rule such as "if there is even one deletion, a person must approve" has to be kept not by a person but by the pipeline.
How it works
The plan a person reads carries symbols.
| Symbol | Meaning | Risk |
|---|---|---|
+ |
Create | Low |
~ |
In-place update | Medium |
- |
Delete | High |
-/+ |
Delete and then recreate | Very high |
+/- |
Create and then delete (create_before_destroy) |
High |
<= |
Data source read | None |
The reason a recreation appears is usually that you touched an argument that cannot be changed, and such lines carry a # forces replacement mark. You must also remember that in the summary line at the very bottom, Plan: X to add, Y to change, Z to destroy, a recreation is counted as 1 on both the add and destroy sides.
What automation needs is not the symbols but two things. One is -detailed-exitcode.
| Exit code | Meaning |
|---|---|
| 0 | No changes |
| 1 | Error |
| 2 | Changes present |
Thanks to these three values, periodic drift detection such as "alert if there are changes" can be made in a single line of shell. The other is the JSON representation. If you convert the plan file to JSON, a resource_changes array comes out, and the change.actions of each entry is ["create"], ["update"], ["delete"], or for a recreation ["delete","create"]. And resource_drift separately holds changes outside the code found in the query stage. That means you can look separately at the changes that will be reflected in the plan and the changes that have already happened.
-target is an option that picks only part of the graph to process. When you use it, the tool issues a warning. Because if you apply only part, the rest can be left diverged from the code and the state can become inconsistent. It is a tool for recovery situations, not a normal workflow.
What you see in the field
First, mass drift after a provider upgrade. When you raise a major version, changes appear on hundreds of resources because of the defaults of newly added attributes. Most of them do not change the actual infrastructure but fill attributes into the state. Both being startled and reverting without telling the two apart, and conversely applying without checking, lead to accidents. If you extract it as JSON and count by action, the judgment is quick.
Second, the cost of detection itself. plan calls the provider API for every managed resource. At large scale it hits API limits, and during detection the state lock is held and conflicts with deployments. The detection interval is not free.
Third, secrets left in the plan log. The plan output prints resource attributes as they are. The habit of pasting it whole into a notification channel or CI log becomes a leak path as it is. It is safer to send only a summary and keep the full output in a place with controlled access.
Three things you must check by eye in a plan
The output of terraform plan is long. You cannot read all of it, so decide what to look for and
then look.
First, the destroy count in the summary line. In Plan: 3 to add, 1 to change, 2 to destroy,
a destroy that is not 0 is always a reason to stop and look. If it is intended, check by
name what is being deleted, and if it is not intended, it is usually a change of a resource address (you
moved a module or changed count to for_each). In that case, instead of delete and create,
move only the address with a moved block.
moved {
from = aws_instance.web[0]
to = aws_instance.web["a"]
}
Second, attributes carrying # forces replacement. A plan in which a database is
replaced just because you changed a name is exposed here. For a stateful resource, putting prevent_destroy
on it blocks it at the plan stage altogether.
lifecycle { prevent_destroy = true }
Third, where (known after apply) is attached. If there are many of these values, it means it is hard to know in advance what the
plan will actually do. It is because they reference attributes of other resources that do not exist yet, and that in itself is normal, but you need to know that, if an
important decision (a security group rule, a policy document) depends on
this value, it cannot be reviewed before applying.
Apply the plan as a file. If someone changes something between the plan and the apply, something different from what you saw gets applied.
terraform plan -out=tfplan
terraform show -json tfplan | jq '[.resource_changes[]
| select(.change.actions | index("delete"))] | map(.address)'
terraform apply tfplan
In automation, use the exit code. -detailed-exitcode returns 0 (no changes),
1 (error), and 2 (changes present). The flow "if there are changes, get a person's approval" in CI
comes from here.
What you will do in the next lab
In /root/tf/plan, you save the plan to a file, convert it to JSON, and count and record by action. You confirm the two exit codes of -detailed-exitcode directly, and make an in-place update and a recreation caught together in one plan. You fix a managed file by hand to see resource_drift, and narrow the scope with -target while reading the tool's warning. Finally you build a review report that judges by itself a plan that includes a deletion.