TT Lab
Get started
Learn Learning paths Courses

FDE Capstone: The Warehouse Got the Same Order Three Times

Health check is green, but orders are not coming in

Continue in TT Lab

In one line

A deployment is finished not when "the new version is up" but when you confirm that "one business transaction goes all the way through on the new version". If you split versions into directories and change them all at once with a single symlink, you can roll back all at once in the same way when the confirmation fails.

Why this was needed

On the day the customer's order intake API was raised to 1.5.0, the deployment owner saw that the health check was green and reported completion. That night the screen showed "received" for every order, but not a single one had gone into the warehouse. The storage path setting was missing in the new version. The health check did not lie. The process was alive and /health returned 200. But that question was "is it alive?", and what the customer wanted to know was "do the orders go in?"

An FDE often works on a customer host with little or no deployment automation. Even on a single VM where you cannot use Kubernetes' rolling updates or a blue-green switch, you have to be able to implement the same principle with a filesystem and a shell script.

How it works

The structure is simple.

/root/site/
├── releases/
│   ├── 1.4.0/        (VERSION, app.py)
│   └── 1.5.0/
├── current -> releases/1.4.0
├── shared/           orders.jsonl · app.pid · app.log  (판이 바뀌어도 남는 것)
└── deploy.log        배포 한 번에 JSON 한 줄

The app always starts from current/app.py. Changing the version is not overwriting files but changing where current points, and rolling back is the same action. The old version's directory is still there, so you do not need to download or unpack anything again to roll back.

The key is "the moment of switching". rename(2) atomically replaces newpath if it already exists, guaranteeing that no other process sees a moment when that name does not exist. If newpath is a symlink, the link itself is overwritten. So you just create a temporary link in the same directory and rename it over current. The reason it has to be the same directory is in the same documentation. A rename between different mounts fails with EXDEV. In Python, os.replace does the same thing, and the documentation says that if it succeeds it is atomic (a POSIX requirement).

Then what about the commonly used ln -sfn? The GNU ln manual explains -f as "remove existing destination files", yet says that unless you use --backup there is no brief moment when the destination is absent. This is not old behavior. The 8.27 (2017-03-08) entry of the coreutils NEWS announces that ln -f A B no longer removes B first. When we measured on the lab image (coreutils 9.4), while one process alternately changed current 3,000 times with ln -sfn and another 3,000 times with a temporary link plus mv -T, another process did readlink about 10 million times, and with both methods it never once saw "absent". So the difference is whether you rely on the tool's version or on the guarantee of the system call. If you do not know which version of ln the customer host has, the rename side is easier to explain.

The one that causes accidents more often is -n. The manual explains -n as "do not treat the last operand specially when it is a symbolic link to a directory". In our measurement, when current pointed to releases/a and you ran ln -sf releases/b current, current stayed the same and a new link called b was created inside releases/a. The command ends in success, so nobody knows.

A health check and a smoke test have different roles.

Check What it asks The method in this lab
Health check Is the new process up and responding? /health 200, and is the version the new version?
Smoke test Does one business path go all the way through? POST one order and then read it back with GET

There is a reason to check the version in the health check. If the old process could not be taken down and the new process dies from a port conflict, the one answering /health is the old version. You have to check first that the green light belongs to the new version. Also, the new version may take more than 2 seconds to load its cache, so you wait up to an upper limit, but if the process is already dead, you do not wait.

What it looks like in the field

The most common mistake in rolling back is rolling back without a record. The next morning the customer asks "I heard 1.5.0 went up last night, so why is it 1.4.0?" You need one line saying what you rolled back, why, and to which version before the conversation can start. You also do not delete the directory of the failed version, because it is something to investigate.

The second is retention cleanup. If, to save disk, you put "keep only the latest 3" in cron, then after a rollback current may not be the latest. If you rolled back 1.6.0 and are back on 1.5.2, and then 1.7.0 and 1.7.1 fail in succession, a cleanup by newest order deletes the version that is running now. You protect current, and the previous version to go back to if current is a problem, regardless of date.

What really matters in practice

What you will do in the next lab

After unpacking a build by hand to make current, you build in turn an atomic switch script, a deploy script with a health check and a smoke test, a rollback, and a retention cleanup. The grader makes, each time with a different version number, builds that are normal, break only in the smoke test, die as soon as they start, and start late, runs your scripts, and measures again the version that is actually responding and the symlink. At the end, you deploy the customer's 1.5.0 yourself, roll it back, and leave an incident record.