TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

Playbooks — Write State, Not Commands

Continue in TT Lab

In one sentence

Each task in a playbook is not "run this command" but a declaration: "when this is done, the state must look like this".

Why this was needed

Anyone who has set up a server with a shell script has seen the script grow longer and longer. At first it was a single line, mkdir /opt/app; on the second run it fails with an "already exists" error, so you switch to mkdir -p; the permissions may differ, so you add chmod; and if they are already right there is no need to run it, so you wrap it in an if. In the end, half of the script becomes code that checks what the current state is.

Ansible's modules shoulder that half for you. Give the file module path, state=directory, and mode=0755, and it creates the directory if it is missing, leaves it alone if it exists, and fixes the permissions if they differ. And it reports changed only when it actually changed something. This report is the foundation for idempotency and handlers, which you will learn in the next modules.

How it works

A playbook is a list of plays, and a play holds "which tasks, in order, on which hosts".

- name: 웹 서버 기본 설정          # 플레이 이름
  hosts: web                      # 대상 (인벤토리의 그룹/호스트/패턴)
  gather_facts: true              # 대상의 정보를 먼저 수집할지
  tasks:
    - name: 앱 디렉터리 준비       # 태스크 이름
      ansible.builtin.file:       # 모듈
        path: /opt/app
        state: directory
        mode: "0755"

There are three rules to remember.

  1. Tasks run from top to bottom, in order. With multiple hosts, however, one task finishes on all hosts before the next task starts.
  2. Give every task a name. Without a name, the log prints the module name and all its arguments, which is hard to read and also hard to select by tag.
  3. command/shell is the last resort. These two do not know the state, so they report changed every time. If a dedicated module exists, use it, and if you have no choice but to use the shell, add a guard such as creates.

There are also ways to check before running. --syntax-check checks the YAML and the play structure, and --check reports only "what would change" without actually changing anything. Adding --diff shows even the differences in file contents. A playbook that runs on a production server for the first time must always go through --check --diff first.

What you see in the field

First, the two faces of tags. With tags, you can push only the configuration again using --tags config, which is convenient. But once the habit of running only part of a playbook by tag sets in, you get "a playbook that has never been run in full". Such a playbook produces an error you have never seen one day when it is finally run in full. Tags are an acceleration device for debugging, not a normal execution path.

Second, the limits of check mode. --check is accurate only when the module supports that mode. A later task that depends on a result produced by command can give nonsense in check mode. So "check mode is clean" and "the real run is safe" are different propositions.

Third, the name is the documentation. When an incident happens, what people read is not the code but the execution log. "nginx install — the config file is overwritten in the next task" is far more useful than "Install nginx".

Where does it stop when it fails

When a playbook fails midway, it leaves a state in which what has been applied and what has not are mixed. Unlike a script, Ansible handles many hosts at the same time, so this state gets more complicated, and if you don't know the default behavior, it is hard to decide what to do after an incident.

By default, only that host drops out. If a task fails on a host, that host is excluded from the later tasks while the remaining hosts keep going. So if only 1 of 10 hosts fails, the other 9 go all the way to the end, and you end up with a cluster split across two versions.

A few mechanisms change this behavior.

The relationship between serial and handlers is especially easy to forget. By default, handlers run once when the play ends, but with serial they run at the end of each batch. So restarts happen batch by batch, and this is in fact the behavior you want in a rolling deployment.

You also need to decide in advance whether it is safe to run again after a failure. If every task is idempotent, you can just run it again, but if even one non-idempotent task is in the middle, rerunning makes things worse. This is where you see why the earlier module said to enforce idempotency with tests. Automation you cannot be sure is safe to rerun is automation that a person has to clean up by hand when it fails.

What you will do in the next lab

Create /root/ans/play/site.yml, express directory creation, config file placement, and editing one line of an existing file as tasks, and use tags and check mode yourself. At the end, fix a playbook with four planted errors and get it to pass.