When the numbers do not add up, the cause is ordering
One-line summary
A race condition is not something that "happens sometimes" but something that always happens once you line up the starting line, and the moment you can reproduce it, it becomes a defect you can fix.
Why this is needed
"Yesterday's settlement was short by 3 records." There is no error in the log. The code looks fine when you read it. It gives the right answer when you run it once. A good share of reports like this are race conditions, and they have one thing in common — run alone, they are always right.
The most common shape is the lost update. A process reads a value, does something with that value, and writes it back incremented by one. If another process read the same value in between, whoever writes later overwrites the earlier result. Two records came in, but only one remains. If ten overlap completely, nine disappear.
The reason the investigation gets stuck here is usually reproduction. The two processes would have to enter that section at exactly the same moment, and when you type by hand, alternating between two terminals, that hardly ever happens. So you build a starting line. If you first bring up all the processes, wait until they are all in a waiting state, and then create a single start marker, from that moment the race is no longer a probability but a reproducible fact.
How it works
The way to prevent it differs with the shape of the problem. If you tell three cases apart, most things fall into place.
1. A read-modify-write section — bind it with a lock. You must make the read and the write into one unit. In Python you take an exclusive lock on the file with flock from the fcntl module. This lock is advisory — a program that opens the file without taking the lock writes without any hindrance. So every program that touches the same file must keep the same promise.
There is a trap here that catches people very often. open(path, "w") empties the file the moment it opens. Even if you take the lock on the next line, it is already too late. Open with "r+" where you read and write, and with "a" where you append, and then take the lock.
2. Something you do after checking — do the check and the action in one step. There is a gap between "create the file if it does not exist" and "take the lease if it is empty." The way to remove this gap is to use the creation itself as the judgment. POSIX open stipulates that if you pass O_CREAT and O_EXCL together, it fails when the file already exists, and that the check and the creation are atomic. In Python it is os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY), and the failure arrives as FileExistsError. The loser receiving the exception is the normal behavior.
3. A file rewritten as a whole — swap it in atomically. If the other side reads while you are writing a configuration file or a state snapshot, it sees half-written content. With JSON that causes a parse error, and that error looks like a bug in the reading side's code. The solution is to write everything to a temporary file in the same directory and then rename it. POSIX rename stipulates that even if the new name already exists, the replacement is atomic, and that at no moment in between does the name disappear. In Python it is os.replace. If you forget the condition that the temporary file must be on the same filesystem, it quietly becomes a copy and the atomicity breaks.
증상 틈이 있는 자리 막는 방법
합계가 모자란다 읽기와 쓰기 사이 배타 잠금으로 한 덩어리
둘 다 자기가 주인이라 한다 확인과 생성 사이 O_CREAT 와 O_EXCL
읽는 쪽이 파싱에 실패한다 쓰는 도중의 파일 임시 파일에 쓰고 rename
파일이 비어 있다 잠금 전에 "w" 로 열기 "r+" 또는 "a" 로 열기
What you see in the field
First, giving up on reproduction and staring at the code. A race condition found by eye is easy to miss, and even if you believe you found one, you cannot know whether it is the cause of that symptom. If you build a starting line and get the symptom in your hands before fixing, the fix is also proved by the same method.
Second, you added a lock and nothing changed. Two common causes are the scope and the target of the lock. If you read outside the lock and only write inside it, it prevents nothing. There are also cases where you lock different files, or where each process locks its own different temporary file.
Third, only one side keeps the courtesy. An advisory lock has meaning only if everyone keeps it. If one operations script overwrites with >, all of that day's locks become meaningless. So the locking convention is kept not by code but by documentation and review.
Fourth, believing it will not happen with one core. A process can be preempted at any time, so regardless of the number of cores, another process can get in between a read and a write. In fact, even on a 2-core Pod, if you start eight processes at the same moment, only one update remains.
What really matters in practice
- Reproduce it by building a starting line. If a race does not reproduce, you cannot even know whether you fixed it.
- Wrap the lock starting from the read. A lock that wraps only the write is decoration.
- Turn check-then-use into a single call. O_EXCL exists for that spot.
- For a file written as a whole, swap it in with rename. The reading side should see only the old file or the new file.
What you will do in the next lab
You receive a harness that lines up the starting line and see with your own eyes that eight lock-free updates leave only one. You do the same thing inside an exclusive lock and see them all remain, and close the check-then-use gap with O_EXCL so that there is always exactly one winner. You swap in a document written as a whole atomically to bring the reading side's parse errors to 0, confirm how the mistake of emptying the file before locking destroys data, and then report the evidence on one page.