Two Runs Must Give the Same Result
In one line
Idempotency is the property that running the same operation several times gives the same result, and without it, automation becomes a gamble whose outcome differs every time it runs.
Why this was needed
Automation is not run just once. A runner dies from a timeout and is rerun, someone else runs the same playbook again, and a scheduler applies the same thing every hour. If "do nothing if it is already done" is not guaranteed at such times, rerunning itself becomes dangerous. In the end people become afraid of automation, and because they are afraid, they do it by hand, and because they do it by hand, they are back to snowflake servers.
The way to prove it is surprisingly simple. Run apply twice in a row and check that the second plan shows no changes (exit code 0). Automation that does not pass this one-line check cannot yet be trusted.
How it works
Idempotency comes from "check the current state before running." Before creating a directory, see whether it exists; before writing a file, see whether the content is the same; before changing permissions, read the current permissions. So a well-built module always has two branches: changed if it changed, and unchanged if it was already right.
The most common place where idempotency breaks is when you run commands directly with shell or command instead of a dedicated module. The tool does not know what that command does, so it runs it every time. That is why you attach conditions such as creates= or removes= so that work that is already done is skipped. The classic bug is to add a line to a configuration file by appending it with >>. Every time it runs, the same line piles up one more, so running it ten times gives ten lines. The correct implementation first checks whether the line already exists, and if there is a different value for the same key, replaces that line in place.
To preview what will change before applying, use check mode and diff together (--check --diff). It shows only what would change without actually changing it, so it is essential the first time you apply to a production server. And the results must be shown as counts. Only with a report that tallies changed and unchanged can a person confirm that changed is 0 on the second run.
What you see in the field
Half of the question "the playbook succeeded, so why is the configuration strange?" comes from non-idempotent tasks. The log printed only ok, but in reality the same line went in three times, or a restart happened every time and the service was being cut off periodically. So mature teams run the playbook twice in CI. If changed on the second run is not 0, the build is broken. They enforce idempotency with a test, not a document.
One more thing: recording state gives you the ability to detect. If you leave the path and hash of each resource at the time of applying, you can later tell just by comparison whether that file was changed from outside. This is exactly what tools with state files do.
Idempotency and convergence are different
The two words are often mixed up, but the scope they guarantee differs. Idempotency is the property that doing the same operation several times gives the same result, and convergence is the property of driving the system to the desired state whatever the current state is. Most automation satisfies only the former and misses the latter.
The difference shows up in "deleting." If you run a playbook that places three configuration files twice, changed is 0. That much is idempotency. But if someone has created a fourth file by hand, this playbook never finds it. That is because the desired state was written only as "these three exist," not "only these three exist." To converge, you have to declare the entire contents of the directory as managed and delete whatever is not on the list.
So when designing automation, you decide two things explicitly. Where does the area I own end, and if I find something unfamiliar inside that area, do I delete it or report it? If you delete it silently, you wipe out someone else's work, and if you do nothing, the servers slowly drift apart.
The same distinction is needed for operations with side effects. A service restart is the typical case. It should restart only when the configuration file has changed, but if you run it unconditionally every time, the playbook itself looks idempotent (because the file contents are always the same), yet in reality the service is cut off every time it runs. This is why the notification (handler) mechanism exists, and if you leave in the result report how many times restarts actually happened, you can close the question "why does it drop for a moment every hour?" in seconds later.
Finally, a note on the relationship with retries. An idempotent operation can simply be rerun when it fails, but for a non-idempotent operation, a retry is a duplicate execution. The key point is that when you fail to get a response because of a network error, you cannot know whether the operation was actually carried out. If you plan to retry, the operation must be idempotent, and if you cannot make it idempotent, there is no choice but to attach a unique marker to each operation and check whether it has already been processed.
What you will do in the next lab
You build four small scripts that ensure directories, files, lines, and permissions, and make each distinguish changed from unchanged exactly. Then you run a runner that bundles the four twice in a row and prove idempotency by whether changed is 0 on the second run. At the end you add a state file, drift detection, and a lock that prevents concurrent runs.