TT Lab
Get started
Learn Learning paths Courses

Debugging in Practice

Eight updates arrived and one survived

Continue in TT Lab

Goal

After lining up the starting line to turn a race condition into a reproducible fact, you prevent the lost update with an exclusive lock, close the check-then-use gap with atomic creation, and swap in a document written as a whole with an atomic rename. Finally, you confirm for yourself what is destroyed by the mistake of emptying the file before taking the lock.

Why it matters

A race condition gets reported as "the numbers are off sometimes." Since run alone it is always right, the investigation gets stuck at reproduction. But if you bring up all the processes first and send them off at the same moment with a single start marker, the race becomes not a probability but something that always happens. From the moment it reproduces, it is a defect you can fix. The way to prevent it differs with the shape of the symptom. You bind a read-modify-write section with a lock, turn check-then-use into a single call, and swap in a file written as a whole with a rename. If you mix these three up, the symptom remains even after you add a lock. The trap that catches people most often is the scope of the lock. open(path, "w") empties the file the moment it opens, so even if you take the lock on the next line, it is already too late. This mistake survives longer because it arises in code that has a lock. The grader does not trust your explanations. It uses its own harness to actually start your workers at the same time and directly counts the final value and the number of winners. The number of processes changes on every run, so you cannot memorize values and plug them in.

Steps

  1. Create and run /root/race/gen_race.py to create /root/race/spawn.py (the starting-line harness).
  2. Send lock-free updates concurrently with /root/race/naive_add.py and write the lost updates in /root/race/lost.json.
  3. Do the same thing inside an exclusive lock with /root/race/locked_add.py and write it in /root/race/lock.json.
  4. Close the check-then-use gap with /root/race/claim.py and write the winner and losers in /root/race/claim.json.
  5. Prevent partial writes with /root/race/safe_write.py and /root/race/read_probe.py and write it in /root/race/atomic.json.
  6. Use /root/race/append_line.py to append without emptying the file before locking, and write it in /root/race/append.json.
  7. Run the two versions several times and leave reproducible evidence in /root/race/evidence.json.
  8. Report in /root/race/summary.json and /root/race/race_report.md in four sections.

Notes

Get the starting-line harness in hand

Create and run /root/race/gen_race.py to create /root/race/spawn.py. This harness starts several workers and then sends them off at the same time with a single marker file.

Where a race condition investigation gets stuck is reproduction. When you type by hand, alternating between two terminals, hardly ever do you get in at the same moment. Save this harness as it is and run it. There are no workers yet, so you build them in the next step.

See eight become one

Create /root/race/naive_add.py (read, work for 300 milliseconds, increment and write), run it concurrently with 8 processes, and write procs, expected, observed, lost, and work_ms in /root/race/lost.json. The counter file starts at 0.

A worker must leave a ready marker (<barrier>.<pid>.up) before waiting, wait until the --barrier file appears, and then start. If you rest for --work-ms between the read and the write, another process reads the same value in the meantime. Do not be surprised by the result — that is the purpose of this step.

Make it one unit with an exclusive lock

Create /root/race/locked_add.py so that everything from the read to the write is inside an exclusive lock, run 8 of it with the same harness, and write procs, expected, observed, lost, and method in /root/race/lock.json. observed must be 8.

Take LOCK_EX on the file with fcntl.flock. If you read the value before taking the lock, it prevents nothing, so be careful about the order. And if you open with "w", the file is emptied before you even get the lock — this is a place to read and write, so open with "r+".

Close the check-then-use gap

Create /root/race/claim.py so that it claims the lease (0 if it wins, 9 if it loses), run 8 of it concurrently, and write procs, winners, losers, owner_in_lease, and method in /root/race/claim.json. There must be exactly one winner.

If you check with os.path.exists and then create, someone else creates it in between. If you pass both O_CREAT and O_EXCL to os.open, the check and the creation happen in one step, and the loser receives FileExistsError. The key point of this approach is that receiving the exception is the normal behavior.

Keep readers from reading a half-written file

Use /root/race/safe_write.py to swap in the document atomically and /root/race/read_probe.py to count while reading, and write rounds, entries, reads, bad_reads, and method in /root/race/atomic.json. bad_reads must be 0, rounds at least 50, and reads at least 100.

If you read while it is being written, JSON parsing fails, and that error looks like a bug in the reading side's code. If you write everything to a temporary file in the same directory and move it with os.replace, the reading side sees only the old file or the new file. If you create the temporary file on a different filesystem, the atomicity breaks.

The mistake of emptying the file before locking

Create /root/race/append_line.py so that it appends its own line without deleting the existing lines, run 8 of it concurrently on a file with a few lines put in beforehand, and write procs, lines_before, lines_after, kept_original, and duplicate_lines in /root/race/append.json.

The trap in this step is not the lock but the way you open. "w" empties the file the moment it opens, so even if you take the lock on the next line, it is already too late. Where you append, open with "a" and then take the lock. lines_after must be lines_before plus the number of processes.

Leave reproducible evidence

Run the lock-free version and the locked version each at least 3 times with the same harness, and write trials, procs, expected, naive_finals, and locked_finals in /root/race/evidence.json. The lengths of the two lists must equal trials.

A single observation may be a coincidence. If you run the same harness several times and write the two lists side by side, you can see that the side without a lock always stays at the same spot and the side with a lock always reaches the expected value. This is the evidence you show the customer.

Report what was lost and what prevented it

In /root/race/summary.json, write procs, expected, naive_observed, locked_observed, claim_winners, bad_reads, and lines_after, and in /root/race/race_report.md, report in four sections: ## 무엇을 잃었나 ## 왜 잃었나 ## 무엇으로 막았나 ## 남은 위험 (in order, these mean: what was lost, why it was lost, what prevented it, and the remaining risk).

What the customer buys is not 'we added a lock' but 'seven of eight records vanished, and now all eight remain.' Write the numbers as they are. Also write the limitation of advisory locks under the remaining risk.