serial and delegate_to: split into batches, hand the work over
Summary in one line
Zero-downtime deployment is not something a deployment tool does for you; it is a matter of how you split up the playbook, and the handles for that are serial, which divides hosts into batches, and delegate_to, which has another machine do part of the work.
Why this is needed
By default, Ansible runs tasks one at a time, in order, across all hosts. It finishes task 1 on every host and then moves on to task 2. If there is a task that restarts a service, all hosts go down at that moment at the same time. With a twenty-host web tier, twenty hosts die together.
That is why you need a limit like "no more than this many at a time." But a limit alone is not enough. If you do not take the server being deployed out of the load balancer beforehand, user requests keep flowing to it, and if you do not check the state after deployment, a broken version spreads to the next batch. Only when these three — splitting, taking out and putting back, and judging each batch — are in place can you call it zero-downtime.
How it works
serial divides the play's host list into batches.
| How to write it | Result for three hosts |
|---|---|
serial: 2 |
Two batches: web1,web2 → web3 |
serial: "50%" |
Two hosts → one host |
serial: [1, "100%"] |
web1 alone first (canary), then all the rest |
Writing it as a list gives a canary deployment. You roll out to one host first, and if it is fine, send the rest all at once. If the last element is smaller than the remaining hosts, it repeats at that size.
There is something you must know here. When you use serial, the play behaves as if it starts over from the beginning for each batch. The proof is that the PLAY [...] header is printed in the execution log as many times as there are batches. As a result, three things become per-batch.
run_once: trueruns once per batch, not once for the whole play.- Handlers fire at the end of the batch, not at the end of the play.
max_fail_percentageis judged per batch.
The last one is the most important. Using serial: 1 and max_fail_percentage: 0 together means "if even one host fails, do not move on to the next host." If the second host fails, the third does not even start, and NO MORE HOSTS LEFT is printed in the log. This is the mechanism that keeps an incident from spreading to three hosts.
any_errors_fatal: true, which looks similar, is a different thing. This one means "if a failure occurs on any host, stop the whole play right there," and it gives no room by percentage. If you want to keep going even when one host out of twenty fails, use max_fail_percentage; if you must stop immediately when even one host fails, use any_errors_fatal.
delegate_to runs a single task on another host. It leaves the target list as it is and moves only the execution.
- name: 로드밸런서에서 뺀다
community.general.haproxy:
state: disabled
host: "{{ inventory_hostname }}"
delegate_to: lb1
There is one point here that is easy to get confused about. Even inside a delegated task, inventory_hostname points to the original host. In the example above, inventory_hostname is not lb1 but web1, which is being deployed right now. That is why "tell the load balancer to take web1 out" can be expressed in one line. In the execution log both names are printed together, as in changed: [web1 -> lb1].
If you also give delegate_facts: true, the facts created by that task are stored under the delegation target instead of the original host. For example, a value looked up once on the load balancer can be seen by all hosts through hostvars['lb1']. If you leave this option out, the facts stay on the original host, and there is nothing in hostvars['lb1'].
The concurrency handles are also worth knowing. forks is a global setting that decides how many hosts the controller handles at once, and serial is the batch size of a play. Of the two, the smaller one becomes the actual number of concurrent executions. throttle is a limit applied to a single task only, so it is used in situations such as "this API allows only so many calls per second." strategy: free makes hosts not wait for each other so the fast hosts finish first, but it does not guarantee order, so it is not suitable for zero-downtime deployment.
What it looks like in the field
First, batch size is a capacity calculation. If you apply serial: "25%" to twenty hosts, five hosts are taken out at the same time. If the remaining fifteen cannot handle normal traffic, the deployment itself is an outage. Batch size should not be a matter of taste but the answer to "how many hosts can be out at once?"
Second, a batch without a health check is meaningless. If you do not check with uri or wait_for after deployment, a broken version goes straight to the next batch. The check task has to fail for max_fail_percentage to do its job.
Third, after taking it out you must put it back. If the play fails during deployment, the last server taken out stays outside the load balancer. So the restore task should go inside block/always, or at the very least you need a separate procedure to check "is there a server that is currently out?"
Fourth, the delegation target must also be in the inventory. delegate_to: lb1 uses that host's connection settings only when lb1 is in the inventory. If it is not, it tries to connect by name alone. That is why you list hosts that "are not targets but are told to do work" — load balancers, bastions, monitoring servers — in the inventory too.
What you will do in the next lab
You list three web hosts and one load balancer in the inventory, split batches with serial, and confirm from the records how many times it actually runs. You build a canary batch as a list, count from the outputs that run_once and handlers run per batch, leave take-out and put-back records on the load balancer side with delegate_to, and use delegate_facts to see under whom facts are recorded. Then you make the middle host fail its health check and confirm that the third batch does not even start, and finally leave a summary file of one full cycle of take out, deploy, check, and restore.