It failed, but the exit code was 0
The retry caused the second outage
Summary
A tool that calls external commands has to prevent two accidents in advance: a command that never finishes and a command that runs twice. A timeout prevents the first; idempotency and locking prevent the second. Retries are safe only after both are in place.
Why this matters
A deployment script called restart-service, and it did not respond. The owner added a "retry". At the next outage, the script called the unresponsive command three times and waited 30 minutes, and in fact all three had run behind the scenes, so the service restarted three times. The retry created a second outage. The problem was not the retry; it was that two things the retry presupposes were missing: a guarantee that the command eventually finishes (a timeout) and a guarantee that running it again gives the same result (idempotency).
How it works
The run() function of subprocess executes a command, waits until it finishes, and returns a CompletedProcess. Three arguments form the skeleton of an operations tool.
timeout=초— when the time is up, it kills the child process and waits for it, then raisesTimeoutExpired. The key is the part the documentation spells out as "killed and waited for". It leaves no zombie.check=True— raisesCalledProcessErrorif the exit code is not 0. The exception carries the arguments, the exit code, and (if captured) stdout and stderr.capture_output=True— captures stdout and stderr. Withtext=Trueyou receive them as strings.
import subprocess
try:
r = subprocess.run(cmd, timeout=30, capture_output=True, text=True)
except subprocess.TimeoutExpired:
return 124 # coreutils timeout 과 같은 코드를 쓰면 셸 사용자가 바로 안다
return r.returncode
A retry has to distinguish the kind of failure. A command that ended with 1 because the network dropped is worth calling again, but a command that ended with 2 because of bad arguments gives the same result a hundred times over. The waiting interval also has to grow: if the other side failed from overload, retrying at a constant interval keeps the overload going. The convention is to grow it exponentially, as in backoff * 2 ** (attempt - 1).
Idempotency is the property that "running once and running many times give the same result". If an external command is not idempotent, the tool puts a check step in front of it: look at a marker file to see whether it has already been applied, and skip if so. Create the marker only after the command succeeds. If a marker appears even though the command failed, the next run believes "it is already done".
If two copies of the same tool run at the same time, the marker check becomes a race. Both see "not yet done" and both run. If you give flock() from fcntl the flags LOCK_EX | LOCK_NB, then when it cannot take the lock it does not wait but raises OSError (the errno is EACCES or EAGAIN; for portability the documentation says to check both). The tool takes that exception to mean "another run is in progress" and backs off. When a process dies, the kernel releases its lock, so stale locks do not get left behind.
There is one more trap. If the command is a shell script and it launched sleep or another command inside, killing only the child (the shell) leaves the grandchild behind. If the grandchild holds the standard output pipe, the pipe never closes and communicate() never returns. If you launch the child with Popen(..., start_new_session=True) as the leader of a new process group and send the signal to the whole group with os.killpg(), the grandchildren are cleaned up too. run() only calls wait() after a timeout, so it sidesteps this hang, but the grandchild stays alive.
The last topic is termination signals. When cron kills a job that ran over time, or when Kubernetes takes down a Pod, the tool receives SIGTERM. If you attach a handler with signal.signal(signal.SIGTERM, handler) from signal, the tool can forward the same signal to its child, wait, and then end with exit code 143 (128 + 15). Without a handler, only the tool dies and the child becomes an orphan and keeps running, which is another form of the accident where the restart command ran three times.
What it looks like in the field
The habit of passing a string to subprocess.run(cmd, shell=True) invites both accidents. If the arguments contain spaces or quotes, it turns into a different command, and when a timeout fires, what dies is the shell, which may not be the real command the shell started. Pass a list and use shell=False (the default). The second is an ordering mistake: "recording success before running the command". Markers, logs, and database updates must come after success. The third is not recording the number of retries. If the log does not show that the third attempt succeeded, nobody knows that the command fails twice every time. If you leave one JSON line per attempt, that log becomes a metric as it is.
What you will do in the next lab
You build the external command runner runner.py. Starting from running a command and returning its exit code unchanged, you add a timeout (124), retries with exponential backoff, an idempotent apply built on marker files, a single-run flock lock, SIGTERM forwarding, a JSON log line for each attempt, and finally --dry-run. The three fixture commands (flaky.sh, which fails the first two times; hang.sh, which never ends; and apply.sh, which takes effect on every call) are in /opt/fixtures/pyops/bin/.