TT Lab
Get started
Learn Learning paths Courses

Infrastructure as Code

What the State File Creates and What It Breaks

Continue in TT Lab

In one line

The state file is the only bridge between the code and the actual infrastructure, and it is also a single point of failure that, if mishandled, can stop the whole team.

Why this was needed

There are four reasons a state file exists. First, mapping. It connects the resource names written in code to the identifiers of the resources actually created. Second, dependency tracking. To know what to delete first and what to create later, it has to remember the relationships. Third, performance. If you query every resource every time, API calls explode, so the last-seen values are used like a cache. Fourth, collaboration. It becomes the reference point for sharing the same facts among team members.

The problem is that this file can hold sensitive information in plain text. A DB password or a certificate goes in as it is. So an encrypted remote backend is practically essential. Attaching sensitive = true to an output is only screen masking, and inside the state file the value is visible as it is. If you confuse the two, you come to the mistaken belief that "it's hidden, so it's safe."

How it works

The mechanism that prevents concurrent applies is the lock. When apply starts, it conditionally creates a lock entry, and if someone else already holds it, a conflict error occurs. What often surprises people here is that the default wait time is 0 seconds, that is, immediate failure. The design is that it is better to fail fast and let a person look at the situation than to hang silently.

Locks can also be left behind. If the network drops, a CI runner dies from a timeout, or someone forces a stop with Ctrl+C, only the lock remains. There is a force-unlock command, but if another user is actually working, the state can be broken, so it is a last resort. There must first be a procedure to check who took that lock and when.

State is handled not by editing it by hand but with dedicated commands. You move with state mv, remove from management with state rm, bring in a resource that already exists with import, and express refactoring with a moved block. In particular, you must know exactly that state rm only removes from management and the actual resource is not deleted. If you mistake it for a delete command, the opposite accident occurs.

When things grow, the state has to be split. If you pile everything into one file, a plan takes more than 10 minutes and API rate limiting kicks in. Split by component, such as networking, compute, DB, and monitoring, to minimize the blast radius. There are also choices in how to separate environments. Workspaces have no code duplication but weak isolation and a wide blast radius. Directory separation creates some duplication but lets you keep a separate backend and IAM for each environment, so the blast radius is narrow. So the basic formula is directory separation for production and workspaces for short-lived test environments.

What you see in the field

Drift divides into three types. Configuration drift, where attributes changed; existence drift, where something was created or deleted outside the code; and dependency drift, where reference relationships broke. The causes are also well known. Manual changes in the console, automatic system changes such as autoscaling, default value changes from a provider upgrade, and parallel applies.

Run detection at least once a day, and attach severity rules. A change to a security group or IAM policy is CRITICAL, and a delete or replace action is always HIGH or above. However, if you try to detect every drift, false positives overflow. Fields managed by an autoscaler, such as desired_capacity or desired_size, must be removed with ignore_changes so that alerts remain signals.

The last is culture. It is unrealistic to completely ban console access during emergency incident response. Instead, establish a culture of syncing the code within 24 hours after an emergency change, and have automatic detection watch over it. It is a design that makes things return rather than blocking them with rules.

When you lose the state file and when you lock it

State is neither the code nor the actual resources but a third truth. The ways in which these three diverge are precisely the list of accidents you run into in IaC.

If the state disappears, the resources remain and only the management vanishes. Terraform sees resources not in the state as "not yet created," so on the next apply it tries to create the same thing again, failing with a name collision, or, if the resource's name is generated automatically, quietly producing two of them. The only recovery is to bring them back one by one with import. That is why for a remote backend you always turn on versioning and deletion protection together.

Secrets go into the state in plain text. If you passed a DB password as a variable, that value is written into the state file as it is. Even if you hide it in the output with sensitive = true, the stored value is the same. This is why the state store must be treated at the same grade as a secret store. For S3, check encryption and access policies; for local, first check .gitignore so it is not committed at all.

Without a lock, two people write different futures at the same time. If two applys overlap, the one that finishes later uploads a state that erases the earlier result. Resources are created but are not in the state, the most troublesome state to fix. The S3 backend locks with a DynamoDB table, and GCS and Azure lock with their own features.

Do not release a leftover lock carelessly. If CI dies midway, the lock remains. Use force-unlock only when really nobody is running. Releasing a running apply amounts to creating the situation of the previous paragraph by hand. First check the person and time written in the lock information.

What you fixed by hand shows up in the next plan. A setting changed in the console in a hurry appears in terraform plan as a plan to revert it. You then have two choices: make the code match reality, or make reality match the code. If you pick the third choice, "just ignore it for now," that resource gets reverted when the next person innocently runs apply.

terraform plan -refresh-only   # 코드는 그대로 두고, 현실과의 차이만 본다
terraform state list           # 상태가 아는 자원의 목록
terraform state show <주소>    # 상태가 기억하는 속성

What you will do in the next lab

With a real declarative tool (OpenTofu), you cause the accidents in this text one after another. You move the state file and see the same server become two, try to overwrite a state with a different lineage and get rejected, get back a lost resource with import, get blocked by the state lock during an apply, and split the state of layers with different lifetimes. In the lab that follows, you find where a secret hidden with sensitive remains in plain text in the state and plan files, and block it with state encryption. In the quiz after the two labs, you check the reason the state file exists, the handling of sensitive information, the dangers of locks and force-unlock, and state splitting and blast radius.