Variable Precedence and Facts
In one sentence
In Ansible, "my variable has no effect" is mostly not because the variable is missing but because a value defined in a stronger place is winning.
Why this was needed
Say you deploy the same application to dev, stage, and prod. Only the port, the replica count, and the log level differ. If you copy the playbook three times, six months later the three files have become different creatures. Pulling values out of the code is not a matter of taste but a structural choice that stops branches from multiplying.
The problem is that there are too many places to put values: a play's vars, a separate vars_files, group_vars/, host_vars/, variables inside the inventory, -e on the command line, a role's defaults and vars, and even values created at run time with set_fact. Each place has a different precedence, and without knowing that order you spend a day on "I clearly changed the value, but it isn't taking effect".
How it works
The detailed order is long, but the skeleton worth remembering in practice is short.
| Strength | Place | Nature |
|---|---|---|
| Weakest | A role's defaults/main.yml |
"This role's default" — a place made to be overridden |
| Weak | Inventory group variables | Values common to an environment |
| Medium | Inventory host variables | An exception for that one server |
| Strong | A play's vars, vars_files |
The value for this run |
| Very strong | Values created with set_fact |
Values computed during the run |
| Strongest | -e on the command line (extra vars) |
Beats everything |
Keeping just two principles solves most cases. First, put values that may be overridden in defaults. Second, use -e only for one-off exceptions. Once -e becomes a habit, the value is recorded nowhere, and the next person cannot reproduce it.
Facts are a different kind of thing. They are not values that people write but facts collected from the target server. With gather_facts: true, the setup module runs at the start of the play, collects things like the OS type, architecture, host name, network interfaces, and memory, and puts them under ansible_facts. With facts, you can automatically take per-target branches such as "apt on Ubuntu, dnf on RHEL".
Collection has a cost. With hundreds of hosts, fact gathering alone takes tens of seconds. A play that does not need facts should turn them off with gather_facts: false, and if you run repeatedly, consider fact caching.
register is the third kind. It stores the entire result of a task (standard output, exit code, whether it changed) in a variable. A common mistake here is writing the whole result object straight into a file. What you need is usually just .stdout.
What you see in the field
First, putting out an urgent fire with -e and not reflecting it in the code. If you saved the service during incident response by passing -e replicas=10, that value must be committed to the repository. Otherwise the next deployment quietly reverts it. This is an incident with exactly the same structure as drift in Terraform.
Second, trust facts but verify them. Facts such as ansible_distribution are mostly accurate, but containers or special images produce unexpected values. When you branch on facts, keep a default branch for unexpected values.
How to check where a value came from
More practical than memorizing precedence is asking directly what that value is on this host right now. Inference can be wrong, but measurement is not.
The fastest way is to print the value with a single task. Run debug ad hoc to print a particular variable, and the final winner for that host comes out immediately. And running it against many hosts at once is especially valuable. If only one machine has a different value, it means an exception is written in that host's host_vars or in its inventory line, and that is usually the cause of the problem.
If you are curious not about the value but about where it came from, narrow the scope step by step. Temporarily delete the group variable, and if the value changes, it came from there; if it does not change, it comes from a stronger place. Also, if you turn on verbose output when running, the log shows which variable files were read, so you can quickly catch the common situation of editing a file that was never read.
Naming also plays a big part in this problem. If you build the habit of prefixing variable names, collisions nearly disappear. Compared with port, myapp_port is better, and a role's variables use the role name as the prefix. A short name can be used by any role, so the moment another role uses the same name, one of the two quietly loses.
Finally, agreeing on a team rule for what goes where is more effective than memorizing precedence. Values that differ by environment go in group variables, an exception for a single server goes in host variables, and defaults that a role proposes go in defaults. If just these three lines are followed, you will hardly ever use the other places, and if you never use them, precedence cannot surprise you.
What you will do in the next lab
You create play variables, a variable file, group variables, and host variables, and see with your own eyes which one wins, and then see command-line -e beat all of them. You collect facts and save them as JSON, and combine values with register and set_fact to build a final summary file.