TT Lab
Get started
Learn Learning paths Courses

CS for Building Good Services — Relearning Textbook Ideas by Measuring

The Gaps Are Between Instructions, awaits and Statements

Continue in TT Lab

In one line

A race condition does not arise "when there are multiple threads"; it arises "when there is a gap between the value you read and the value you write where someone else can cut in." With threads, that gap is between bytecode instructions; with asyncio, it is at every await; and with a database, it is between SELECT and UPDATE. You have to either remove the gap (an atomic operation) or let only one task at a time pass through the whole gap (a lock).

Why this was needed

The concurrency module of the "Operating Systems" course has you build, by hand, the scene where four threads increment the same value and lose updates, and a deadlock. Even people who finish that lab often get three things wrong in service code. Python has a GIL, so x += 1 is safe. asyncio has only one thread, so there are no races. The database will block it for me. All three are wrong, and this module confirms all three with numbers.

How it works

What the GIL guarantees and what it does not. The glossary defines the GIL as the mechanism that makes CPython execute bytecode in only one thread at a time. The Library FAQ explains that Python usually switches threads only between bytecode instructions, so a single instruction is atomic, and says that L.append(x) is atomic but i = i+1 and D[x] = D[x] + 1 are not. If you use dis to disassemble counter += 1, on this lab image (3.12) you get four instructions: LOAD_GLOBAL, LOAD_CONST, BINARY_OP, and STORE_GLOBAL. The read and the write are different instructions, so if the thread switches in between, the value someone else incremented gets overwritten. How often it switches is decided by sys.setswitchinterval, but the documentation says that value is only an "ideal" time slice and the operating system decides which thread takes over. That is why races rarely show up in tests and appear when load rises. The lab inserts time.sleep(0) between the read and the write to widen the gap on purpose. It does not create a bug; it is a device that makes the bug that was already there visible every time.

A lock has to wrap the whole gap. The Lock in threading is acquired when entering a with statement and released when leaving it. A common half-fix is to wrap only the write in a lock. The value you read is already stale, so even if you line up the writes, stale values are simply recorded one after another. Check-then-act, such as seat reservation where you "check whether it is free, then reserve," is the same: if you wrap only the check in a lock, two requests come out holding the answer "free."

In asyncio, await is the gap. The event loop runs on one thread, and tasks hand over control only at an await. A stretch with no await is effectively atomic, and a stretch with an await in it is not. If you check the balance, await to record, and then deduct, during that await another withdrawal from the same account checks the same balance. The asyncio synchronization primitives documentation recommends using asyncio.Lock with async with, and says acquisition is fair, so the coroutine that waited first enters first. The same documentation states firmly that these tools are not thread-safe, and that for OS thread synchronization you should use threading. Conversely, if you await while holding a threading.Lock inside a coroutine, the next task waits for that lock and stalls the loop thread itself.

Calls that block the loop. The asyncio development guide explains that if you call a function that uses the CPU for 1 second directly, all tasks and I/O are delayed by 1 second, and says debug mode logs callbacks that take more than 100ms. time.sleep or a synchronous driver call stalls the loop in exactly the same way. The lab measures how late a task that wakes every 10ms actually woke compared with schedule, as "loop lag." asyncio.to_thread runs a function on another thread and frees the loop, but as the documentation says, because of the GIL it is usually effective only for I/O-bound functions.

In a database, the gap is between statements. SQLite's transaction documentation says that if a write comes during a read transaction, it is upgraded to a write transaction only when possible, and if another connection has already changed or is changing the data, it fails with SQLITE_BUSY (locking stages). Conversely, if you read with SELECT without a transaction, compute, and write with UPDATE, you lose the update with no error at all. Python's sqlite3 module, under its default settings, opens a transaction implicitly only before INSERT, UPDATE, DELETE, and REPLACE, so a SELECT is not protected. UPDATE t SET n = n + 1 puts the read and the write in one statement and removes the gap.

Place Gap How to block it
Threads Between bytecode instructions threading.Lock around the whole read-to-write
asyncio await asyncio.Lock around the whole check-to-deduct
DB Between statements An atomic UPDATE, or a transaction that surfaces failure + retry

The GIL and speed. The threading documentation says that in CPython only one thread at a time executes Python code, so to use multiple cores you should use multiprocessing or ProcessPoolExecutor, and that threads are still appropriate for running several I/O tasks concurrently. The glossary says the GIL is always released during I/O. Since 3.13 you can build a free-threaded build with the GIL turned off, but it is not the default build, and that guide also recommends using threading.Lock rather than relying on the internal locks of built-in types. It means that, with or without the GIL, the code has to close the gap of read-modify-write.

What it looks like in the field

Node has the same structure. The "It Wasn't One Request - Everything Got Slow" course covers event loop delay through observation and alerting, so the loop lag measurement in this module is the Python version of that. Row locks, SELECT FOR UPDATE, and deadlocks are covered by the "Concurrency — When Two People Touch One Row" course, and using a profiler to see that adding threads does not speed up CPU work is covered by the "Diagnosing CPU and Memory Leaks" course. This module is the place in between where you measure where the gaps are.

What you will do in the next lab

You will disassemble counter += 1 into bytecode, lose updates with threads, and then block it with a lock. You will create and fix a double seat booking and an overdraft between awaits, then measure how long a blocking call stalls the loop and move it with to_thread. You will compare a silent loss and a loud failure with two SQLite connections, and finally measure the speed of threads and processes under the GIL.