It failed, but the exit code was 0
So the retry does not cause a second outage
Goal
Make a Python tool that calls external commands prevent "a command that never finishes" and "a command that runs twice": timeouts, retries, idempotency, locking, signal handling, and per-attempt logs.
Why it matters
Retries are safe only on two premises: the command eventually finishes, and running it again is safe. A retry added without those premises triples the outage. The timeout of subprocess.run kills the child and waits for it, the marker file is left only after success, flock blocks concurrent runs, and a SIGTERM handler leaves no orphan processes. This lab puts all four into one small runner.
The fixture commands are in /opt/fixtures/pyops/bin/: flaky.sh (fails the first two times and succeeds the third; the attempt count is kept in the file named by the environment variable FLAKY_STATE), hang.sh (does not finish for 60 seconds), and apply.sh <이름> (on every call it appends one line to APPLY_LOG, by default /root/pyops/sub/applied.log; it fails if APPLY_FAIL=1). Here the placeholder is the name to apply.
Steps
- Create
/root/pyops/sub/runner.py.runner.py run -- <명령...>runs the command with subprocess.run, writes the command's standard output unchanged to its own standard output, and ends with the command's exit code. Here the placeholder stands for the command and its arguments. - Add
--timeout <초>, where the placeholder is a number of seconds. When the time is up, the command is killed, one line containingtimeoutis printed to standard error, and the tool ends with exit code 124. No command process may be left behind. - Add
--retries <n> --backoff <초>, where the second placeholder is a number of seconds. If the exit code is not 0, wait backoff × 2^(attempt-1) seconds and try again, up to n times. A timeout (124) also counts as a failure and is retried. - Create
runner.py apply <이름>, where the placeholder is the name to apply. If the marker/root/pyops/sub/state/<이름>.doneexists, printalready applied, end with 0, and do not call apply.sh. If it does not exist, call/opt/fixtures/pyops/bin/apply.sh <이름>and create the marker only if it succeeds (0). On failure, end with 1 without a marker. applytakes an fcntl.flock(LOCK_EX | LOCK_NB) lock on/root/pyops/sub/state/lock. If it cannot take the lock, printanother run in progressto standard error and end with exit code 3.- When it receives SIGTERM, forward SIGTERM to the running command, wait for it to finish, and then end with exit code 143. No command process may be left behind.
- For every attempt, append one JSON line to
/root/pyops/sub/runs.jsonl. The keys arets(an ISO 8601 string),cmd(a list of strings),attempt(starting at 1),rc(an integer), andduration_ms(an integer). apply --dry-run <이름>, where the placeholder is again the name, does not call apply.sh and does not create a marker; it prints to standard output what it would do, including the name, and ends with 0. If the name has already been applied, it printsalready applied.
Notes
- subprocess.run(cmd, timeout=...) kills the child and waits for it before raising TimeoutExpired. If you use Popen directly, launch it with
start_new_session=Trueand send the signal to the whole group withos.killpg(os.getpgid(proc.pid), sig). In hang.sh the shell launchessleep 60, so if you kill only the shell, the grandchild sleep stays behind holding the pipe. - flock is taken on an open file descriptor. Keep the lock file open until the run ends. The kernel releases the lock if the process dies.
- To test:
FLAKY_STATE=/tmp/f1 python3 runner.py run --retries 3 --backoff 0.2 -- /opt/fixtures/pyops/bin/flaky.sh - Common mistakes: passing a string with shell=True, creating the marker before running the command, and waiting the same interval between retries.
Run the command and return its exit code unchanged
Create /root/pyops/sub/runner.py. runner.py run -- <명령...> runs the command with subprocess.run, writes its standard output unchanged, and ends with the command's exit code. Here the placeholder stands for the command and its arguments.
Create run with argparse subparsers and accept command with nargs=argparse.REMAINDER. Pass a list and do not use shell=True. Returning r.returncode is enough.
End the command that never finishes
Add --timeout <초>, where the placeholder is a number of seconds. When the time is up, the command is killed, one line containing timeout appears on standard error, and the tool ends with exit code 124. No command process may be left behind.
subprocess.run(timeout=...) kills the child and waits for it before raising TimeoutExpired. If you use Popen, launch it with start_new_session=True and kill the process group with os.killpg. If the grandchild sleep of hang.sh holds the pipe, communicate() does not return when you kill only the child.
Retry with exponential backoff
Add --retries <n> --backoff <초>, where the second placeholder is a number of seconds. On failure, wait backoff × 2^(attempt-1) seconds and try again, up to n times. A timeout also counts as a failure.
Loop over range(1, retries + 2) and return immediately if rc == 0. The waiting time doubles with each attempt. Delete the FLAKY_STATE file and test with flaky.sh.
An idempotent apply with a marker
If /root/pyops/sub/state/.done exists, runner.py apply <이름> prints already applied and ends with 0. Otherwise it calls apply.sh and creates the marker only if that succeeds. On failure it ends with 1 without a marker.
Call write_text for the marker only after the command ends with 0. If you create it before running, a failed change stays recorded as applied. Test the failure path with APPLY_FAIL=1.
Block concurrent runs with a lock
apply takes fcntl.flock(LOCK_EX | LOCK_NB) on /root/pyops/sub/state/lock. If it cannot, it prints another run in progress to standard error and ends with exit code 3.
Keep the lock file open with open('w') and call fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB). If the errno of the OSError is EACCES or EAGAIN, another run is holding it.
Forward SIGTERM to the child
When the runner receives SIGTERM, it sends SIGTERM to the running command, waits for it to finish, and then ends with exit code 143. No command process may be left behind.
Attach a handler with signal.signal(signal.SIGTERM, handler) and call send_signal(signal.SIGTERM) on the Popen object from inside the handler. When communicate() returns after the handler ran, return 143.
One JSON line per attempt
For every attempt, append one JSON line to /root/pyops/sub/runs.jsonl. The keys are ts, cmd (a list of strings), attempt (starting at 1), rc (an integer), and duration_ms (an integer).
Write json.dumps(dict) + '\n' in append mode. A timeout (124) is also an attempt, so record it. Multiply the time.monotonic() difference by 1000 and make it an integer.
Say what it would do without doing it
apply --dry-run <이름>, where the placeholder is the name, does not call apply.sh and does not create a marker; it prints a plan containing the name to standard output and ends with 0. If the name has already been applied, it prints already applied.
You may handle dry-run before the lock rather than after it, because it changes nothing. Look only at whether the marker exists and choose one of the two sentences.