Lifecycle and Drift — Which Way Do You Converge
In one sentence
The key question in responding to drift is not "how do I fix it" but "do I make the code match the real thing, or make the real thing match the code?"
Why this was needed
At three in the morning, traffic surges and the service dies. The person on call goes into the console, bumps up the instance type, raises the autoscaling maximum, and briefly opens a debugging port. The service comes back. Up to here, these are right judgments.
The problem is what comes after. These three changes were not recorded anywhere. The code still says the old values, and so does the state file. A few days later, when someone runs an apply for an entirely different reason, the tool faithfully reverts all of the early-morning changes. The outage recurs, and nobody knows the cause.
This is drift. A mismatch among three things: the code (the state that should exist), the state file (the state last known), and the real thing (the state that exists now). Drift itself is not a sin. Undetected drift is the sin.
How it works
The basic detection tool is plan. A plan refreshes the real thing before running, compares it with the state, and then compares the state with the code. That is enough for people to read but not enough for automation. An approach that parses the output strings breaks when the version changes.
That is why -detailed-exitcode exists.
| Exit code | Meaning |
|---|---|
| 0 | No changes — code, state, and real thing match |
| 1 | Error |
| 2 | Changes present — drift or unapplied changes |
With only these three values, you can attach drift detection to cron or CI as is. A trap that often catches people here is set -e. It treats exit code 2 as a failure and the script dies, so the verdict command must always be run with exception handling.
To look more precisely, save the plan with plan -refresh-only and then extract it with show -json. The JSON has a separate resource_drift array, so you can distinguish differences that came from code changes from differences that arose outside the code. If you mix "someone touched it in the console" and "there is code not yet deployed" in the same alert, people soon start ignoring alerts.
The lifecycle block is four knobs that intervene in this flow.
| Argument | What it does | Where to use it |
|---|---|---|
create_before_destroy |
Creates the new one first and then deletes the old one | When you must avoid interruption during replacement |
prevent_destroy |
Blocks a destroy plan itself with an error | State stores, production DBs |
ignore_changes |
Excludes differences in specific attributes from the plan | Attributes managed by an external system |
replace_triggered_by |
Recreates this resource when another resource changes | Forced replacement of resources with no update mechanism |
ignore_changes is especially important. If you catch as drift even values that intentionally change outside the code, such as a replica count managed by an autoscaler or tags the console attaches automatically, false alarms overflow, and overflowing alarms soon become ignored alarms. However, the ignore list must contain exactly those attributes only. If you ignore wholesale, real incidents go quiet along with them.
prevent_destroy stops with an error at the plan stage. It means that when you delete a resource block by mistake or rename it, a brake is applied before the destruction is approved. In exchange, when you delete on purpose you must first remove this line from the code, so it is also a device that makes deletion always go through review.
What it looks like in the field
First, there are three response directions. (1) Revert the real thing to match the code with apply, (2) if the change was intended, fix the code to match the real thing, (3) if the real thing is the right answer and the code is already correct, update only the state with apply -refresh-only. If the early-morning emergency action was a right judgment, reverting it is itself an incident. Running automatic recovery without judgment is the most dangerous.
Second, automate by severity. A single mismatched tag may be reverted automatically, but a security group rule or an instance type must be looked at by a person. The standard pattern in practice is selective recovery, which automatically recovers only low severity and stops with an alert for high severity.
Third, detection itself has a cost. A plan calls the provider API for every managed resource, so if you run it often on large-scale infrastructure, you hit API limits. During detection a state read lock is also held, which can conflict with a deployment running at the same time. And plan output can include sensitive values such as passwords, so you must filter it when saving it to logs.
Fourth, mass drift after a provider upgrade. If a new version changes attribute defaults, hundreds of drift items show up even though nobody did anything. In that case it is the provider that changed, not the real thing, so pinning versions and checking the changelog come first.
What to do in the next lab
You actually apply create_before_destroy and prevent_destroy and confirm that deletion is blocked, and with ignore_changes you see the plan stay empty even when you change the code value. After making another resource's change trigger recreation with replace_triggered_by, you create drift by hand-editing the real thing without going through code. You detect it with a -refresh-only plan and exit code 2, and respond in opposite directions, once by reverting to the code and once by absorbing the real thing into the code. At the end, you build a script that spits out the detection result as JSON to give automation its shape.