TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

Tags and partial runs: a sliced run promises no final state

Continue in TT Lab

In one sentence

A tag is a device for choosing "only this one this time", and it is inherited downward from the place where it is attached. But a run cut by tags does not guarantee the final state the playbook promises — knowing that fact is more important than the tag syntax.

Why this was needed

Playbooks grow. At first there were five tasks, and half a year later there are eighty. One day comes when you want to fix just one line of a configuration file. Running everything takes fifteen minutes, from fact gathering to package checks, and in the meantime twenty tasks you did not want to touch go by.

Tags are the answer to this problem. You put labels on tasks and run only those carrying a label with --tags config. Conversely, you can leave out only the slow ones with --skip-tags slow. The syntax is simple, and so people start using it in production the day after they learn it.

This is where incidents happen. A tag is a knife that cuts the run, but a playbook is not designed to be cut. The directory that task 8 created is used by task 12, and the handler receives the configuration changed by task 20 and makes the service reread it. If you run only task 12 with --tags config, it fails because the directory is missing, or worse, it pretends to succeed while leaving a half-correct state.

So the subject of this module has two layers. The tag syntax and inheritance are one layer, and how to recognize a playbook that can be cut is the other.

How it works

Tags attach in four places and are inherited downward.

Where it attaches Inheritance scope
Task That one task
Block Every task inside the block
Play Every task of that play
Role / import_tasks / include_tasks Explained below

If you run --list-tasks, the inheritance becomes visible. If you put playtag on a play, playtag follows onto the tag list of every task. If you put blocky on a block, each of the two tasks in the block gets blocky.

import_tasks and include_tasks diverge here. import_tasks is static — the tasks are expanded in place when the playbook is read, and the tag attached to the import statement is inherited by each expanded task. include_tasks is dynamic — tasks are inserted only during execution. So a tag attached to an include statement attaches only to the include statement itself and is not inherited inside.

The result of this difference is bewildering at first sight. If you give --tags included, the include statement runs because its tag matches, but the tasks inserted inside it have no included tag, so they are all filtered out. Nothing happens, with no error at all. The official documentation advises that if you want to tag the tasks inside an include too, use the apply keyword or use import_tasks instead.

There are five special tags.

Tag Meaning
always Runs always, unless you explicitly exclude it with --skip-tags always
never Never runs unless you call it by name directly
tagged Every task that has at least one tag
untagged Every task that has no tag at all
all Everything (the default)

Put always on tasks that must not be missing from any selection, such as collecting inventory facts or setting common variables. never lets you write down dangerous tasks — deleting a database or redeploying everything — in the playbook while keeping them from running by a slip of the hand. On a task that has never, if you attach a label such as danger together with it, it runs only when you call it with --tags danger.

There is one spot to watch out for. If you put a tag on a play, every task of that play becomes "tagged". Then --tags untagged selects nothing, and --tags tagged selects everything. Inheritance thus even changes the meaning of the special tags.

Look at the list before running. --list-tags shows what tags this playbook has, and --list-tasks shows, in order, what would run for this selection. Neither connects to the target or changes anything. If you give them together with --tags, you get a list that reflects that selection as it is — the habit of looking at this first when you use tags in production for the first time halves incidents.

There is also a way to run from the middle. --start-at-task "태스크 이름" starts from the task with that name. You use it so as not to rerun from task 1 when a long playbook died at task 40. --step proceeds by asking a person at every task — it is interactive, so it cannot be used in automation, but it is useful when you follow a playbook you are seeing for the first time by hand.

--start-at-task is riskier than tags. This is because it starts from the middle in a state where everything the earlier tasks made and every notify the earlier tasks sent is absent. Even if failed=0 comes out at the end, the system is not in the state the playbook promised.

To sum up the two ways a partial run is dangerous:

First, dependent tasks are left out. Task 7 uses a value that the set_fact of task 3 set, and if you choose only task 7 with --tags, that variable stays undefined. If you are lucky, it fails immediately, and if you are unlucky, a default puts in something unexpected.

Second, handlers do not run. A handler runs only if it receives a notify. If the task that changes the configuration is left out of the selection, nobody calls the handler and it quietly passes. Conversely, in a run that chose only the configuration task, the handler runs normally — so whether a handler runs is decided by "who notified", not by "tags". If you confuse the two, you will look in the wrong place for "I applied a tag and the service was not restarted".

What you see in the field

First, doing the first deployment with --tags. If you apply only --tags config to a new server, it fails because the directory is missing. Tags are a tool for a system where the whole playbook has already run once. Do the first convergence with everything.

Second, tags start to replace documentation. When there are thirty tags, nobody knows which one is safe to apply. Keep tags to around five, and give names, in a script or a Makefile, to the combinations you use as a set.

Third, an incident because never was not applied. Some teams keep the "full reinstall" task commented out. A comment eventually gets uncommented. With a never tag, it remains as code but does not run by accident.

Fourth, trusting the green light after running from the middle. Even if a deployment revived with --start-at-task ends in success, what the earlier tasks should have left behind is not there. After a recovery, the real exit condition is to run the whole thing once more and confirm changed=0.

Fifth, the honest limits of this lab environment. --step is an interactive mode that proceeds only when a person presses a key, so a grader cannot judge it and it was left out of the lab. Tagging roles is also not covered here because roles themselves are the subject of ansible-advanced.

References

What you will do in the next lab

You put tags on tasks and narrow with --tags, look first at what would run before it runs with --list-tags and --list-tasks --tags, and leave some out with --skip-tags. You put tags on a block and a play and check how the inheritance shows up in the list, apply always and never and measure for yourself what the special tags mean, and put import_tasks and include_tasks side by side to separate, by list, the side where tags are inherited from the side where they are not. After running from the middle with --start-at-task, you confirm that the configuration stays at the old value and the handler did not run, and finally you build a tool that judges which tag selections can run to completion.