Blue-Green Switchover and Canary Analysis
This lab runs on a real VM
This box is not a Pod but a virtual machine launched by KubeVirt. A separate Linux kernel runs in it, systemd actually manages services, and docker is a real Docker engine, not an imitation. A container started with docker run becomes a real process, and both docker exec and docker logs work as usual.
This lab used to run inside a Pod. That box had dropped every kernel capability, so the step that starts a container was blocked, and you learned by working around it and unpacking image archives by hand. The workaround is no longer needed.
There are two things to know.
- The first start takes a little over a minute. The VM boots and installs Docker, so it is slower than a Pod lab (usually 40 seconds).
- There is no browser preview. Only one grading port is open for connections into the VM. If you start a web server, check it with
curlfrom inside the VM.
Goal
You build a blue-green environment with two containers, and implement traffic switching, a health check, automatic rollback, canary percentage distribution and an automatic abort verdict in shell.
Why it matters
The difference between deployment strategies is ultimately an exchange between rollback speed and traffic control precision. Blue-green brings up two environments at the same time, so resources cost twice the Pods, but if a problem shows, you can revert the traffic wholesale and instantly. Canary exposes precisely in 1% units, but needs steps when reverting. So canary is used for most services, and blue-green for services where even a single error is costly, such as payments or authentication. In this lab the active target is expressed as one line in a file. Who changes that one line and when, what is checked after changing it, and how it is reverted if the check fails — that is all there is to deployment automation.
Steps
- Start a container whose response body contains
version=bluewith the nameci3-blueand publish it on127.0.0.1:8091. That string must show up incurl http://127.0.0.1:8091/. - In the same way, start
ci3-greenon127.0.0.1:8092with the bodyversion=green. At that timeci3-bluemust keep running. - Create
/root/ci3/switch.sh <blue|green>. It writes the active target name to/root/ci3/active, but if theACTIVE_FILEenvironment variable is set, it writes to that path. For a value other thanblue/green, it writes nothing and exits with a non-zero code. When you finish this step, the contents of/root/ci3/activemust begreen. - Create
/root/ci3/health.sh <URL>. If the target responds normally it gives exit code 0, and if there is nobody there or it fails, a non-zero code. Grading checks with 8091, 8092, and 8099 where nobody is. - Create
/root/ci3/deploy.sh <대상>. It remembers the current active value, switches to the target and then health-checks the target's port. On failure, it reverts the active value to the previous value and ends with a non-zero code. On success, the active value stays as the target and the exit code is 0. The ports must be overridable with thePORT_BLUE(default 8091) andPORT_GREEN(default 8092) environment variables, andACTIVE_FILEmust also be respected. - Create
/root/ci3/canary.sh <퍼센트>. It sends 100 requests, counts where they went and printsblue=<수> green=<수>. The sum of the two numbers is always 100, and if the argument is 0 it must beblue=100 green=0, and if 100,blue=0 green=100. - Create
/root/ci3/analyze.sh <지표파일> <임계퍼센트>. The metrics file hastotal=1000anderrors=10, one line each. The error rate iserrors * 100 / total, and if it is at or below the threshold the exit code is 0, and if it exceeds, it prints the actual error rate and gives a non-zero code. Iftotal=0, there are no samples, so do not let it pass but treat it as a failure. - Create
/root/ci3/deploy-report.json. The fields arestrategy(blue-greenorcanary),active(must equal the current contents of/root/ci3/active),previous(must differ from active),rollback_used,error_rate_pctandabort_threshold_pct. Inrollback_used, write whether a rollback actually happened in this deployment as a JSON boolean (trueorfalse) — the string"true"is not allowed.error_rate_pctmust be at or belowabort_threshold_pct.
Notes
- Example of creating the body: in
/root/ci3/blue/index.htmlputversion=blue, and rundocker run -d --name ci3-blue -p 127.0.0.1:8091:80 -v /root/ci3/blue:/usr/share/nginx/html:ro nginx:1.27-alpine. - You cannot bind ports below 1024. Always publish a high port like 8091/8092 on 127.0.0.1.
- For
rollback_used, you just write the truth, whethertrueorfalse. However, it must be a JSON boolean, and if you wrap it in quotes like"false", it fails the type check. - If you use a sequence number (for example, the remainder when divided by 100) instead of random numbers for the canary distribution, the result is reproducible.
- Common mistakes: taking blue down while bringing green up, switch.sh ignoring
ACTIVE_FILEand always writing to a fixed path, and recording active and previous as the same value.
Bring up the blue environment
Start a container whose response body contains version=blue with the name ci3-blue and publish it on 127.0.0.1:8091. That string must show up in curl http://127.0.0.1:8091/.
The container name is exactly ci3-blue and the published port is 127.0.0.1:8091. You cannot bind ports below 1024. The response body must contain version=blue, so the simplest way is to create an index.html and mount it.
Bring up the green environment side by side
In the same way, start ci3-green on 127.0.0.1:8092 with the body version=green. At that time ci3-blue must keep running.
Start ci3-green on 8092 with the body version=green. You must not take blue down at this time. Both environments must be alive at the same time for a zero-downtime switch to hold.
The active target switch script
Create /root/ci3/switch.sh <blue|green>. It writes the active target name to /root/ci3/active, but if the ACTIVE_FILE environment variable is set, it writes to that path. For a value other than blue/green, it writes nothing and exits with a non-zero code. When you finish this step, the contents of /root/ci3/active must be green.
It is switch.sh <blue|green> and the default record file is /root/ci3/active. If the ACTIVE_FILE environment variable exists, you must use that path (grading verifies with a temporary path). For an unknown value, write nothing and fail. One typo sends the traffic to a place that does not exist.
Health check
Create /root/ci3/health.sh <URL>. If the target responds normally it gives exit code 0, and if there is nobody there or it fails, a non-zero code. Grading checks with 8091, 8092, and 8099 where nobody is.
health.sh <URL> gives 0 if alive and a non-zero code otherwise. Call it in a form that errors on failure and has a timeout, like curl -fsS -m 3. A health check that always passes is the same as having none.
Roll back automatically on failure
Create /root/ci3/deploy.sh <대상>. It remembers the current active value, switches to the target and then health-checks the target's port. On failure, it reverts the active value to the previous value and ends with a non-zero code. On success, the active value stays as the target and the exit code is 0. The ports must be overridable with the PORT_BLUE (default 8091) and PORT_GREEN (default 8092) environment variables, and ACTIVE_FILE must also be respected.
deploy.sh <대상> must remember the current active value first so that it can roll back. The ports must be overridable with PORT_BLUE (default 8091) and PORT_GREEN (default 8092), and ACTIVE_FILE must also be respected. On a health failure, restore the previous value + a non-zero exit code.
Percentage distribution
Create /root/ci3/canary.sh <퍼센트>. It sends 100 requests, counts where they went and prints blue=<수> green=<수>. The sum of the two numbers is always 100, and if the argument is 0 it must be blue=100 green=0, and if 100, blue=0 green=100.
canary.sh <퍼센트> sends 100 requests and prints blue=<수> green=<수>. The sum must always be 100, so you must not just let request failures slip by. Sequence-based distribution is more reproducible than random numbers.
Automatic abort verdict
Create /root/ci3/analyze.sh <지표파일> <임계퍼센트>. The metrics file has total=1000 and errors=10, one line each. The error rate is errors * 100 / total, and if it is at or below the threshold the exit code is 0, and if it exceeds, it prints the actual error rate and gives a non-zero code. If total=0, there are no samples, so do not let it pass but treat it as a failure.
It is analyze.sh <지표파일> <임계퍼센트> and the metrics file is the two lines total=1000 and errors=10. If equal to the threshold, it passes. If total is 0, there are no samples and you cannot judge, so you must not let it pass. When aborting, print the actual error rate.
Deployment report
Create /root/ci3/deploy-report.json. The fields are strategy (blue-green or canary), active (must equal the current contents of /root/ci3/active), previous (must differ from active), rollback_used, error_rate_pct and abort_threshold_pct. In rollback_used, write whether a rollback actually happened in this deployment as a JSON boolean (true or false) — the string "true" is not allowed. error_rate_pct must be at or below abort_threshold_pct.
The active in /root/ci3/deploy-report.json must equal the contents of /root/ci3/active, and previous must differ. In rollback_used, write true/false for whether a rollback actually happened in this deployment — it must be a JSON boolean without quotes. error_rate_pct must be at or below abort_threshold_pct.