TT Lab
Get started
Learn Learning paths Courses

Terraform in Practice

State Surgery — Fixing the Memory Without Touching Reality

Continue in TT Lab

In one sentence

state mv and state rm do not change the infrastructure. They change only the tool's memory. If you misunderstand this one sentence, incidents happen, and if you understand it, refactoring becomes safe.

Why this was needed

Infrastructure code inevitably goes through refactoring. Because you do not like a resource name, because you want to bundle some resources into a module, because the state file has grown too large and you want to split it. But the moment you rename something in the code, the tool judges that the resource with the old name has vanished and a resource with the new name has appeared. The plan shows destruction and creation side by side. If it were a production DB, it is over the moment you approve that plan.

There is a problem in the opposite direction too. During incident response, someone created a security group in the console. The real thing exists, but it is not in the code or the state. The next apply does not know about that resource and so does not touch it, but it means one more resource that nobody manages. Bringing this resource under code management is what import does.

These two directions — moving addresses and bringing in things from outside — are all there is to state surgery.

How it works

The state file is a ledger that records "the addresses of the resources I created and the attributes I last knew." There are four surgical tools.

Command What it does Effect on the real thing
state list Prints the list of addresses None
state show <주소> Prints all attributes of that address (the placeholder is the address) None
state mv A B Moves entry A in the ledger to the name B None
state rm A Deletes entry A from the ledger None (the real thing stays as it is)
import Registers the real thing in the ledger None

The most often misunderstood one is state rm. This is not a delete command but a declaration of giving up management. The real thing stays and only the tool forgets. So if, after state rm, you do not also delete the corresponding block from the code, the next plan says, "that resource doesn't exist, let's create it anew." On a real cloud, it fails trying to create another resource with an overlapping name, or in the worst case a duplicate resource is created. If you want to delete even the real thing, you must use destroy.

What does the same job as state mv in code is the moved block. The difference is in the record. state mv is run once in someone's terminal and vanishes, but moved remains in the code so that the same migration happens automatically when teammates just run plan. So the default is state mv for state used alone, moved for code used by many.

import also has two forms. The command-line tofu import <주소> <ID> (the placeholders are the address and the ID) runs immediately and changes the state. An import {} block is declared in the code, checked first with the plan, and then reflected with apply. The latter can be reviewed and leaves a history, so for collaboration the block approach is better. Both share the point that you must create the place to bring it into (the resource block) in advance. Without an empty shell, the tool does not know where to put that ID.

What it looks like in the field

First, a copy before surgery is not negotiable. Copy the file before every command that touches the state. With a remote backend, pull it down with state pull. A wrongly moved address is hard to undo, but with a copy you can simply revert. The way to check that the copy is genuine is to compare the lineage value — it must be a file with the lineage of the same state.

Second, state corruption usually comes from an interrupted apply. If an apply dies with the network cut, the file breaks. The standard recovery procedure is to restore the previous version using the remote backend's versioning and realign with the real thing using apply -refresh-only. This is the real reason running on a single local file is risky.

Third, always read the first plan after an import. If the real thing's actual attributes differ from the code, the plan tries to eliminate that difference. If you apply before copying a console-created resource's detailed settings into the code, the settings are gone the moment you bring it in. It is also a common defense to temporarily put prevent_destroy on right after an import.

Fourth, think of state-moving work together with locking. On a remote state shared by many, if someone else runs an apply during surgery, the result cannot be predicted. Announce the work time and, if possible, pause the pipeline for a moment.

What you must do before touching state

state rm, state mv, and import do not touch actual resources, but they change the eyes through which Terraform sees the world. If you get it wrong, a live resource leaves management, or the next apply deletes it.

First, take down the state. Even with a remote backend, leave a copy locally.

terraform state pull > state.backup.json
terraform state list | tee resources.txt

Whether there is a way back decides whether you may touch it.

import puts it only into the state. Without code, the next plan says "I will delete it." So the order is write the code first, import, and check that the plan is empty. If the plan is not empty, it means the code differs from the actual resource, so fix the code to match.

terraform plan -generate-config-out=generated.tf   # 코드 초안을 만들어 준다
terraform import aws_s3_bucket.logs my-logs-bucket
terraform plan     # 여기서 "No changes" 가 나와야 끝난 것이다

The Korean comments in this code block say, in order, that the first command generates a draft of the code, and that "No changes" must appear at the last plan for you to be done.

If you write an import block in the code, you can preview it in the plan, which is safer.

import {
  to = aws_s3_bucket.logs
  id = "my-logs-bucket"
}

state rm is not throwing the resource away but letting it go. The actual resource stays and only Terraform forgets. It is used when moving to another stack, and since until you import it at the destination, it is in a state nobody manages, make sure nobody applies in between.

When you move a module, the addresses change wholesale. Instead of state mv-ing resources one by one, if you use a moved block, it shows up in the plan and can be reviewed. State manipulation leaves no record, but a move written in code remains in a commit.

Before touching, check whether there is a lock. If you fix the state while CI is running, the two sides push different states.

What to do in the next lab

You first make a copy and look into the state, then rename a resource with state mv and confirm the plan is empty. With the same command you move a resource into a module, and prove with the plan the fact that state rm does not delete the real thing. Next, you bring outside resources into the state in two ways, with an import {} block and with the tofu import command, and at the end you prove with the exit code of -detailed-exitcode that the code and state match completely.