TT Lab
Get started
Learn Learning paths Courses

Ansible in Practice

Block it before you run it: syntax, lint and preconditions in one gate

Continue in TT Lab

Goal

Starting from a bad playbook that the syntax check lets through, you raise it layer by layer with ansible-lint's rules and profiles, learn to place exceptions in the narrowest scope, block preconditions first with assert, and bundle these three into a single gate script.

Why it matters

The dangerous thing about Ansible is that even a badly written playbook runs fine. A task with no name, a shell command containing a pipe, a file write that sets no permissions all finish green. The trouble comes six months later — reports are always changed so nobody reads them, file permissions differ from server to server, and the name of the failed task is shell so even looking at the log you cannot tell what died. A linter pulls those six months forward to right before the commit. But the moment you turn on a linter, the team immediately meets the next problem — hundreds of findings come out, and a few of them really do need exceptions. If you cannot tell apart turning a rule off in the whole repository and removing just one line at that point, the linter becomes an empty shell within a month. This lab sets up that distinction together with the place that blocks, with assert, what the linter cannot see (whether values make sense).

Steps

  1. Set the default inventory in /root/anslint/ansible.cfg to ./inventory/hosts.ini, and list web1 and web2 of the web group in that file (both with ansible_host=127.0.0.1 ansible_port=2222, and ansible_user in [all:vars] is root). In /root/anslint/messy.yml, write one unnamed play with three tasks — a shell task containing a pipe and no name, a command: mkdir -p task whose name starts with a lowercase letter, and a task written on one line like copy: content=... dest=.... Run ansible-playbook --syntax-check messy.yml and save the output to /root/anslint/out/syntax.txt.
  2. Save the output of ansible-lint -f pep8 messy.yml to /root/anslint/out/lint_before.txt (a format with one finding per line). Then lint the same file in JSON format and save only the rule ids that were hit, without duplicates, one per line in alphabetical order, to /root/anslint/out/rules.txt.
  3. Write /root/anslint/site.yml from scratch — one named play with three named tasks. The first creates the /root/anslint/out/data directory, the second writes the single line port=8080 to /root/anslint/out/app.conf, and the third writes that host's name in uppercase on one line to /root/anslint/out/upper-<호스트이름>.txt (with the host name in the placeholder). Do not use any shell command. ansible-lint --profile basic site.yml must pass, and run the playbook for real and save the output to /root/anslint/out/run.txt.
  4. Change every module in /root/anslint/site.yml to its FQCN (ansible.builtin.<모듈>, with the module name in the placeholder), and add mode to every task that creates a file or directory. ansible-lint --profile production site.yml must pass, and save its output to /root/anslint/out/lint_production.txt.
  5. Add a fourth task to /root/anslint/site.yml — run tar -czf /root/anslint/out/bundle.tgz -C /root/anslint/out app.conf with ansible.builtin.command and attach changed_when: false. The linter catches this task as command-instead-of-module, but the unarchive module only unpacks and cannot bundle, so a shell command is right here. Add a comment that removes just that one line from the rule, make --profile production pass again, and run the playbook again to actually create the bundle.
  6. Add a fifth task to /root/anslint/site.yml — its name is nginx health probe, starting with a lowercase letter, and it writes the single line ok to /root/anslint/out/health.txt. Then create /root/anslint/.ansible-lint and write profile: production, messy.yml and out/ under exclude_paths, and name[casing] under skip_list. Run ansible-lint with no arguments to confirm that the whole directory passes, and save the output to /root/anslint/out/lint_repo.txt.
  7. Create /root/anslint/checks.yml — on localhost it does not gather facts, and sets app_port: 8080, app_env: staging, and allowed_envs: [staging, prod] as play variables. The first task checks that app_port is an integer and between 1024 and 65535 inclusive, and the second task checks that app_env is in allowed_envs. Write both with ansible.builtin.assert and attach fail_msg and success_msg. Save the output of a run with the defaults as they are to /root/anslint/out/assert_ok.txt, and the output of a run overridden with -e app_port=80 to /root/anslint/out/assert_fail.txt.
  8. Create /root/anslint/gate.sh — it moves into the directory given as the first argument (the default is the current directory) and checks three things in order. For each *.yml directly under that directory, it does a syntax check and prints OK syntax <파일> or FAIL syntax <파일> (with the file name in the placeholder); it runs ansible-lint with no arguments and prints OK lint or FAIL lint; and if checks.yml exists, it runs it and prints OK assert or FAIL assert. If even one fails, it exits with a non-zero value. Run this gate in the current directory and save the output to /root/anslint/out/gate.txt.

Notes

A bad playbook that passes the syntax check

Set the default inventory in /root/anslint/ansible.cfg to ./inventory/hosts.ini, and list web1 and web2 of the web group in that file (both with ansible_host=127.0.0.1 ansible_port=2222, and ansible_user in [all:vars] is root). In /root/anslint/messy.yml, write one unnamed play with three tasks — a shell task containing a pipe and no name, a command: mkdir -p task whose name starts with a lowercase letter, and a task written on one line like copy: content=... dest=.... Run ansible-playbook --syntax-check messy.yml and save the output to /root/anslint/out/syntax.txt.

--syntax-check reads the YAML and looks only as far as whether the play and task structure makes sense and whether module names exist. It does not look beyond that — a command that is not idempotent, a task with no name, and a file write with no permissions all pass. The purpose of this step is to see with your own eyes why you must not trust the syntax check. Write it badly on purpose.

Count what the linter catches by rule id

Save the output of ansible-lint -f pep8 messy.yml to /root/anslint/out/lint_before.txt (a format with one finding per line). Then lint the same file in JSON format and save only the rule ids that were hit, without duplicates, one per line in alphabetical order, to /root/anslint/out/rules.txt.

You choose ansible-lint's output format with -f — there is the default format for people to read, pep8 with one finding per line, and json for tools to read. In each JSON entry, the rule id is under the name check_name. You can extract it with jq and sort -u. Rule ids separate details with square brackets, as in name[play] — even for the same name rule, the id tells you which spot was hit. The linter exits with a non-zero value when it finds violations, so keep that in mind when you save.

Raise it up to the basic profile

Write /root/anslint/site.yml from scratch — one named play with three named tasks. The first creates the /root/anslint/out/data directory, the second writes the single line port=8080 to /root/anslint/out/app.conf, and the third writes that host's name in uppercase on one line to /root/anslint/out/upper-<호스트이름>.txt (with the host name in the placeholder). Do not use any shell command. ansible-lint --profile basic site.yml must pass, and run the playbook for real and save the output to /root/anslint/out/run.txt.

A profile is a layer that bundles rules — in the order min · basic · moderate · safety · shared · production, each one up is stricter, and a higher profile includes all the rules of the lower profiles. Do not try to go all the way to production at once; raising one layer at a time is the order used in practice. What basic mostly catches is "there is no name" and "written in one-line free form." If you change a shell command into a module, no-changed-when disappears along with it.

Two more things the production profile requires

Change every module in /root/anslint/site.yml to its FQCN (ansible.builtin.<모듈>, with the module name in the placeholder), and add mode to every task that creates a file or directory. ansible-lint --profile production site.yml must pass, and save its output to /root/anslint/out/lint_production.txt.

Of the things the production profile additionally requires, these two are the ones you run into most often. FQCN prevents name collisions — as collections increase, a name like copy appears in several places, and a short name can become a different module depending on the search path order. The rule that tells you to write mode prevents permissions being left to luck. If you do not write it, the target's umask decides, and that value differs per server. Which rule belongs to which profile appears in the summary table at the end of the lint output.

Keep the rule alive and remove only one line

Add a fourth task to /root/anslint/site.yml — run tar -czf /root/anslint/out/bundle.tgz -C /root/anslint/out app.conf with ansible.builtin.command and attach changed_when: false. The linter catches this task as command-instead-of-module, but the unarchive module only unpacks and cannot bundle, so a shell command is right here. Add a comment that removes just that one line from the rule, make --profile production pass again, and run the playbook again to actually create the bundle.

If you attach # noqa: <규칙id> (with the rule id in the placeholder) to any line of a task, only that task is removed from that rule. To remove several rules, list them separated by spaces. This and skip_list in the configuration file are entirely different in nature — noqa means "an exception only here," so the person next to you can ask why in review, while skip_list means "this rule is not looked at in the whole repository," so that rule effectively disappears. When you place an exception, always choose the narrowest scope. The rule id that got hit now has the same form as the list you extracted in step 2.

Write rules that apply to the whole repository in the configuration file

Add a fifth task to /root/anslint/site.yml — its name is nginx health probe, starting with a lowercase letter, and it writes the single line ok to /root/anslint/out/health.txt. Then create /root/anslint/.ansible-lint and write profile: production, messy.yml and out/ under exclude_paths, and name[casing] under skip_list. Run ansible-lint with no arguments to confirm that the whole directory passes, and save the output to /root/anslint/out/lint_repo.txt.

With a configuration file, you do not have to give --profile by hand every time, and CI and people's hands run with the same rules — this is the real reason to have a configuration file. exclude_paths means "do not look at this path at all." You put in bad examples left for teaching or code made by others. However, if you give a file name directly as an argument, the exclusion list is ignored — exclusion is a rule for "when scanning." What goes into skip_list must be an exception the team has agreed on. Here, you can assume it was decided to use task names that start with a lowercase product name.

Make the playbook ask itself whether values make sense

Create /root/anslint/checks.yml — on localhost it does not gather facts, and sets app_port: 8080, app_env: staging, and allowed_envs: [staging, prod] as play variables. The first task checks that app_port is an integer and between 1024 and 65535 inclusive, and the second task checks that app_env is in allowed_envs. Write both with ansible.builtin.assert and attach fail_msg and success_msg. Save the output of a run with the defaults as they are to /root/anslint/out/assert_ok.txt, and the output of a run overridden with -e app_port=80 to /root/anslint/out/assert_fail.txt.

assert checks whether all the conditions written in that are true. If even one is false, the play stops on that host — that is the purpose. Preconditions must be checked before changing anything. If you stop after deploying about half of it, undoing it is much more expensive. If you do not write fail_msg, the failure message comes out as the raw condition expression and the recipient does not know what to fix. A value passed with -e is a string unless otherwise specified — here you will see why is integer is false. A failing run exits with a non-zero value, so keep that in mind when you save the output.

Bundle the three layers into one gate

Create /root/anslint/gate.sh — it moves into the directory given as the first argument (the default is the current directory) and checks three things in order. For each *.yml directly under that directory, it does a syntax check and prints OK syntax <파일> or FAIL syntax <파일> (with the file name in the placeholder); it runs ansible-lint with no arguments and prints OK lint or FAIL lint; and if checks.yml exists, it runs it and prints OK assert or FAIL assert. If even one fails, it exits with a non-zero value. Run this gate in the current directory and save the output to /root/anslint/out/gate.txt.

The value of a gate lies in blocking. A script that only lets things through and always exits with 0 is the same as having none, and is actually worse because it creates the illusion that "checking is happening." So after you make it, you must feed it bad input — put one deliberately bad playbook in a temporary directory and give that directory as the argument. There is also a reason for putting the three layers in this order. A syntax check takes less than 1 second, lint takes a few seconds, and the precondition check actually runs Ansible — you run the cheap ones first so that a bad commit gets sent back quickly.