Modules — Boundaries Before Reuse
In one sentence
A module is not a device for making code shorter but a device for limiting the radius over which a change spreads.
Why this was needed
Say you build the same stack in dev, stage, and prod. At first, copying the directory three times is the fastest way. It really is fast. The trouble comes six months later. One setting rushed into prod is not in dev, and is half in stage. The three directories have become different creatures by now, and to find the cause of a "problem that happens only in prod" you have to put the three files side by side and run a diff.
The real cost here is not the number of duplicated lines. It is the fact that you have to make a change three times, and that if you miss one of the three, nobody knows. A module flips this problem around to "one implementation, three sets of inputs." Then a change happens in one place, and differences between environments show up as values, not as code.
How it works
A module is just a directory that contains .tf files. There is no special syntax, only conventions.
| File | Role |
|---|---|
variables.tf |
The inputs this module accepts. The module's public API |
main.tf |
The implementation. The part that should not be visible from outside |
outputs.tf |
The values sent out. The module's return values |
The reason for splitting into three files is not taste but reading order. When you look at someone else's module for the first time, what you want to know is "what goes in and what comes out," and the implementation comes after.
There are four properties to remember.
First, a module is a capsule. Resources inside a module cannot be referenced directly from outside. An address like module.primary.local_file.box.filename does not exist. To get a value out, it must pass through an output. It looks annoying, but thanks to this constraint, renaming a resource inside a module does not break the side that uses it.
Second, addresses change. The state address of a resource inside a module gets a prefix of the form module.<호출이름>. (where the placeholder is the name of the module call). When nested, it keeps stacking, as in module.bundle.module.left.local_file.box. The fact that refactoring is an address change connects to the state surgery in lesson 3.
Third, the root injects providers. If you put a provider block inside a child module, the side that uses that module cannot change the provider configuration, and worse, you can no longer delete that module call from the code. Destroying resources needs the provider configuration, but that configuration is inside the module that has vanished. So put only required_providers in the module and the actual configuration in the root.
Fourth, call the same source multiple times. If you copy the module directory because you want more instances, you are back to the original copy-paste problem. The right answer is to call the same source under different names, like module "primary" and module "secondary".
What it looks like in the field
First, the criterion for splitting is "do they change together?" If you split by resource type (all the vpc in a vpc module, all the iam in an iam module), it looks tidy on the surface, but real changes always span several modules. Things that are born together and die together should be bundled into one module. The criterion for splitting state files is the same — it is blast radius and apply time, not folder aesthetics.
Second, validation is cheapest at the module boundary. If a bad value gets inside the module, the error comes from a resource far below. If you block it at the entrance with the validation block of a variable, the error message prints "which variable of which module was rejected and why" as is. Debugging time differs by multiples.
Third, modules have versions. A local relative path is convenient for learning and for a single repository, but for a module used by several teams you need a registry and version pinning. Otherwise, someone else's single commit changes our prod plan.
When to create a module, and how much to hide
You learn that modules are made for reuse, but in practice the right time to make one is when a second place to use it appears. A module used in only one place just makes you open one more file.
A module made too early is sure to be abstracted wrongly. That is because you built it knowing only the first use's needs. When a second use appears, you add one variable, and at a third you add another, until it becomes a module with forty variables. By that point the module hides nothing, so it has no reason to exist.
A good module holds decisions. It should hold not a bundle of resources but the judgment "in our organization, this kind of thing is made this way." Decisions such as turning encryption on, where to send logs, how to attach tags, and how many days to keep backups go inside, and only a few are exposed outside.
Fewer variables are better, and variables that only pass through are bad. If there are many variables that send internal resource attributes straight to the outside, the module is a thin shell. In that case it is easier to read if you remove the module and use the resources directly.
Pin versions and bump them. Fix the version with a registry or a git tag, and have the users state that version explicitly. If you leave it at ref=main, the moment you modify the module, a change nobody planned spreads to every user.
module "db" {
source = "git::https://git.example.com/tf-modules.git//rds?ref=v2.3.0"
...
}
Nest at most two levels. Once it becomes a module in a module in a module, nobody can trace which variable is passed down how far. The code that passes values only grows, and when you read the plan, resource addresses get long and you cannot see what is changing.
Outputs are a contract. Users depend on those outputs, so if you rename one, every user breaks. Export only what is needed, and name it by its meaning from the outside, not by the inner implementation.
What to do in the next lab
You build a small child module called filebox with three files for inputs, implementation, and outputs, and call the same source under two names, primary and secondary. You see for yourself how the state addresses of resources inside the module change, and build a nested structure in which a module calls another module. At the end, you reject a bad value at the module boundary with variable validation, and gather the outputs of three modules into one map.