Roll out with batches and load-balancer delegation, without downtime
Goal
You learn to write a playbook that splits a set of hosts into batches, takes the server being deployed out of the load balancer and puts it back, and judges the state of each batch so that an incident does not spread to the next batch.
Why it matters
By default, Ansible finishes one task on all hosts before moving to the next task. If there is a task that restarts a service, all hosts go down together at that moment. That is why zero-downtime deployment is built not with a tool but with the structure of the playbook. You split (serial), take out and put back (delegate_to), and judge each batch (max_fail_percentage). If even one of these is missing, it is not zero-downtime — if you do not take hosts out, requests go to the server being deployed, and if you do not judge, a broken version spreads to the remaining servers. And when you apply serial, the meaning of run_once and handlers quietly changes too. If you use them without knowing this, "a task you thought would run only once" runs as many times as there are batches.
Steps
- In
/root/ans/roll/inventory/hosts.ini, list a web group (web1, web2, web3,ansible_host=127.0.0.1,ansible_port=2222) and an lb group (lb1,ansible_connection=local). In/root/ans/roll/p01.yml, applyserial: 2and have each host write the members of its own batch, joined with commas, to/root/ans/roll/out/01-<호스트>.txt(with the host name in place of the placeholder). - In
/root/ans/roll/p02.yml, giveserialas the list[1, 100%]to make a canary batch, and write the batch members to/root/ans/roll/out/02-<호스트>.txtin the same way as in step 1. - In
/root/ans/roll/p03.yml, applyserial: 2, and with a single task carryingrun_once: trueanddelegate_to: lb1, leave a line in/root/ans/roll/out/03-runonce.txtin the formatbatch <배치구성원> ran on <실행한호스트>(the placeholders stand for the batch members and the host that ran it). - Make
/root/ans/roll/p04.ymlwithserial: 1, delegate every task to lb1 to create the/root/ans/roll/out/lb/directory, and then for each host leavedrained-<호스트>.txtandenabled-<호스트>.txt, each with the contenthost=<호스트>(the host name goes in the placeholder). - In
/root/ans/roll/p05.yml, use aset_factwith bothdelegate_to: lb1anddelegate_facts: trueto setlb_pool_sizeto the number of hosts in the play, read that value throughhostvars, and leave it in/root/ans/roll/out/05-lbfact.txtin the formatlb_pool_size=<값> on_web=<웹 호스트에서 본 값>(the placeholders stand for the value and the value as seen from a web host). There must be no value on the web host side. - Make
/root/ans/roll/p06.ymlwithserial: 2, and have the task that places/root/ans/roll/out/06-<호스트>.conffor each host (the host name goes in the placeholder) notify a handler. The handler, withrun_once: trueanddelegate_to: lb1, leaves a linerestarted: <배치구성원>in/root/ans/roll/out/06-handlers.txt(the placeholder stands for the batch members). - Make
/root/ans/roll/p07.ymlwithserial: 1andmax_fail_percentage: 0, and sethealthy_hostsin the playvarsto[web1, web3]. There are two tasks: a deployment marker that leaves/root/ans/roll/out/07-deployed-<호스트>.txt(the host name goes in the placeholder), and anassertcheck that looks at whether the host is inhealthy_hosts. Save the execution output to/root/ans/roll/out/07-run.txt. - Put two plays in
/root/ans/roll/p08.yml. The first play runs web withserial: [1, 100%]andmax_fail_percentage: 0, leaves a batch record in/root/ans/roll/out/08-batches.txt, and for each host leavesdrained-,release-(contentrelease=2.4.0), andenabled-markers under/root/ans/roll/out/08/, checking the state withassertin between (all three hosts are healthy). The second play collects those records on lb1 and writesbatches,drained, andenabledto/root/ans/roll/out/rolling.json. Finally, run the same playbook once more and save the output to/root/ans/roll/out/08-run2.txt, and it must showchanged=0.
Notes
- web1, web2, and web3 imitate the one sshd inside this Pod (127.0.0.1:2222) under three names, and lb1 is a local connection. Whether execution really went to a different machine cannot be proven in this environment, so you check it through the records the delegation leaves and the
web1 -> lb1display in the execution log. - You count how many times the batches ran by the number of
PLAY [...]headers in the execution log. throttleandstrategy: freeshow up only in timing and concurrency, and that measurement is unstable in this environment, so they were left out of the lab. See the official documentation section in the reading.- Common mistake: taking hosts out but forgetting to put them back. If the play fails midway, the last server taken out stays outside the load balancer.
- Common mistake: editing the output files by hand. Grading looks not only at the files but also re-runs the playbook in check mode to see whether the same results still come out.
- Strategies and more · Delegation and local actions · Error handling · Playbook keywords
Roll out two hosts at a time
In /root/ans/roll/inventory/hosts.ini, list a web group (web1, web2, web3, ansible_host=127.0.0.1, ansible_port=2222) and an lb group (lb1, ansible_connection=local). In /root/ans/roll/p01.yml, apply serial: 2 and have each host write the members of its own batch, joined with commas, to /root/ans/roll/out/01-<호스트>.txt (with the host name in place of the placeholder).
Who is in the current batch can be found during execution through a variable. If you write ansible_play_batch joined with commas, how the batches were divided is left in the file.
Canary one host first, then the rest at once
In /root/ans/roll/p02.yml, give serial as the list [1, 100%] to make a canary batch, and write the batch members to /root/ans/roll/out/02-<호스트>.txt in the same way as in step 1.
The elements of the list can mix numbers and percentages. Putting 100% as the last element means "all the rest."
run_once is once per batch
In /root/ans/roll/p03.yml, apply serial: 2, and with a single task carrying run_once: true and delegate_to: lb1, leave a line in /root/ans/roll/out/03-runonce.txt in the format batch <배치구성원> ran on <실행한호스트> (the placeholders stand for the batch members and the host that ran it).
Even if the same line is run several times, the file must not grow. Use a module that adds a line only if it is not already there. Counting how many lines come out reveals the scope of run_once.
Take it out of the load balancer and put it back
Make /root/ans/roll/p04.yml with serial: 1, delegate every task to lb1 to create the /root/ans/roll/out/lb/ directory, and then for each host leave drained-<호스트>.txt and enabled-<호스트>.txt, each with the content host=<호스트> (the host name goes in the placeholder).
Even inside a delegated task, the target is the original host. If you use that name as it is in the file name and content, "who was taken out" is left in the record on the load balancer side. In this step it must happen per host, not per batch.
Under whom to record the facts
In /root/ans/roll/p05.yml, use a set_fact with both delegate_to: lb1 and delegate_facts: true to set lb_pool_size to the number of hosts in the play, read that value through hostvars, and leave it in /root/ans/roll/out/05-lbfact.txt in the format lb_pool_size=<값> on_web=<웹 호스트에서 본 값> (the placeholders stand for the value and the value as seen from a web host). There must be no value on the web host side.
The number of hosts in the play can be found during execution through a variable. When you read the same name on the web host side it may be absent, so give a default so that 없음 (meaning "none") is printed.
Handlers run at the end of the batch, not the end of the play
Make /root/ans/roll/p06.yml with serial: 2, and have the task that places /root/ans/roll/out/06-<호스트>.conf for each host (the host name goes in the placeholder) notify a handler. The handler, with run_once: true and delegate_to: lb1, leaves a line restarted: <배치구성원> in /root/ans/roll/out/06-handlers.txt (the placeholder stands for the batch members).
Handlers are linked by name. Counting how many lines are left tells you when the handler runs — if it ran once at the end of the play, there should be one line.
If one host collapses, stop there
Make /root/ans/roll/p07.yml with serial: 1 and max_fail_percentage: 0, and set healthy_hosts in the play vars to [web1, web3]. There are two tasks: a deployment marker that leaves /root/ans/roll/out/07-deployed-<호스트>.txt (the host name goes in the placeholder), and an assert check that looks at whether the host is in healthy_hosts. Save the execution output to /root/ans/roll/out/07-run.txt.
This playbook fails on purpose. When you capture the output to a file, do not let the script stop because of the failure. Whether the deployment marker of the third host appears or not is the answer to this step.
Go through one full cycle and leave a summary
Put two plays in /root/ans/roll/p08.yml. The first play runs web with serial: [1, 100%] and max_fail_percentage: 0, leaves a batch record in /root/ans/roll/out/08-batches.txt, and for each host leaves drained-, release- (content release=2.4.0), and enabled- markers under /root/ans/roll/out/08/, checking the state with assert in between (all three hosts are healthy). The second play collects those records on lb1 and writes batches, drained, and enabled to /root/ans/roll/out/rolling.json. Finally, run the same playbook once more and save the output to /root/ans/roll/out/08-run2.txt, and it must show changed=0.
Delegate taking out and putting back to the load balancer, and do the deployment on the target. If you count the markers with a module that finds files and read the batch record with a module that reads files in, the second play gets shorter. For the second run to be quiet, every task must be idempotent.