Somebody Changed It in the Console
In one line
Drift is when the state declared by the code and the actual state have diverged. Most of the problems that organizations adopting IaC actually wrestle with are here.
Why this was needed
An outage happened on Friday night. Someone went into the console and added one rule to a security group, putting out the urgent fire. It was the right call.
The problem is Monday. When someone else runs plan to make an unrelated change,
it says it will delete the rule that was added on Friday. Because the code
does not have that rule. If you apply without knowing, the outage recurs.
This is drift. And it is not a defect of the tool but expected behavior. A declarative tool takes the code as the truth and makes the actual match it.
How drift arises
| Path | Example | Response |
|---|---|---|
| Emergency manual change | Console operation during incident response | Put reflecting it in code afterward into the procedure |
| Another team's change | The security team adjusts a policy | Make ownership boundaries clear |
| Autoscaling | The instance count keeps changing | Exclude with ignore_changes |
| Cloud defaults | Tags automatically assigned at creation | Reflect in code or ignore |
| Another codebase | Two repositories manage the same resource | Strictly forbidden — there is one owner |
The last item is the most dangerous. If two state files manage the same resource, they keep reverting each other. Each resource must have exactly one owner.
Detection — plan is a diagnostic tool
plan is not something to run only right before a deployment. Running it periodically
to detect drift is good operations.
# 아무 변경도 안 했는데 plan 에 diff 가 있다 = 드리프트
plan → 변경 없음 ✅ 코드와 실제가 일치
plan → 변경 있음 ⚠️ 조사 필요
Run plan daily in CI and send an alert if the result is not empty.
That way you learn of Friday's change on Saturday morning rather than on Monday.
Resolution — three options
When you find drift, the answer is one of three.
1. Absorb it into the code — if the change was right, reflect it in the code. It is the most common and usually the right choice.
2. Revert it — if the change was wrong, force the code's state with apply.
However, you must first check why such a change was made.
3. Take it out of management — for a value that changes by nature, like autoscaling,
exclude only that attribute with ignore_changes. The knack is to narrow it
to the attribute level, rather than removing the whole resource.
Import — bringing what already exists into the code
Adopting IaC does not usually start from a blank slate. There are already hundreds of resources made by hand, and you have to move them into code.
The one thing you must never do here: delete and recreate. For a resource in operation, that is downtime, and if it has data, that is data loss.
Instead, use import. It registers the actual resource in the state file so that the tool recognizes it as "this is something I manage."
1. 자원의 실제 설정을 조사한다
2. 그와 똑같은 코드를 작성한다
3. import 로 상태에 등록한다
4. plan 을 돌려 '변경 없음' 이 나올 때까지 코드를 다듬는다
Step 4 is the key. The plan must be empty for the import to be finished. If you move on with a diff remaining, that resource will change on the next apply.
When handling the state file
The state file holds the actual IDs, attributes, and dependency relationships of resources. Credentials sometimes go in as plain text as well.
- Keep it in a remote backend and turn on the lock. If two people apply at the same time, the state is broken.
- Turn on versioning. It is the only means of undoing a wrong apply.
- Do not commit it to git. Secrets may be in it.
- Do not edit it by hand. If you really must, use the state commands the tool provides.
What you see in the field
- A new hire ran
applyfor the first time and deleted someone else's manual change → they did not read the plan. - The autoscaling group size shows up as a diff every time → there is no
ignore_changes. - The state file is local and two people overwrite each other → no remote backend and no lock.
What to look at in the next reading
In the reading that follows, we look at where to place module boundaries when replicating the drift response across several environments. Connect which decisions among absorb, revert, and exclude belong in a common module and which must stay with the calling environment to reduce the blast radius of changes.