CS for Building Good Services — Relearning Textbook Ideas by Measuring
Reference Counting, the Cycle Collector and the Growing Line
In one line
CPython clears most objects on the spot with reference counting, and only the cycles that reference counting cannot clear are cleared in occasional batches by the generational collector. When that "occasionally" lands in the middle of a request, the average stays the same but p99 jumps. A problem of growing memory has to be looked for not at "the biggest line" but at "the line that grew the most."
Why this was needed
The "Diagnosing CPU and Memory Leaks" course catches leaks from outside the process by the slope of RSS and Private_Dirty. That method tells you "it is leaking," but not "which line." The "The heap had room, but the service stopped" course reads the stalled time from the JVM's GC logs. A Python service has the same two questions: what percentage of the latency tail is due to GC, and which line of the source the growing memory comes from. Both can be answered with numbers using only the standard library.
How it works
Reference counting and the cycle collector. The gc module documentation says this collector supplements the reference counting that Python already uses, so you can turn it off if you are sure you do not create reference cycles. A tree in which a parent holds a list of children and the children point back at the parent keeps holding each other even after the request ends, so the count never reaches 0. Only such cycles are the collector's job. The CPython internal documentation explains that the collector tracks only container objects, which can hold other objects.
Generations and thresholds. The same gc documentation divides objects into three generations, and says that when the number of allocations minus the number of deallocations since the last collection exceeds threshold0, it inspects starting from generation 0, and when generation 0 has been inspected more than threshold1 times, it looks at generation 1 as well. As the internal documentation explains, the oldest generation does a full collection only when the newly added share of long-lived objects exceeds 25%. The default threshold values differ between versions. Printing gc.get_threshold() on this lab image (3.12.3) gives (700, 10, 10), and the latest version of the internal documentation gives the initial value of the default build as (2000, 10, 10). So do not memorize the numbers; print them. The key point is that collection is triggered not by "time" but by "the number of allocations." A handler that creates cycles piles up allocations without deallocations, so it crosses the threshold often.
How to measure collection time. If you put a function into gc.callbacks, it is called with phase "start" just before a collection and "stop" just after, and info contains the generation collected and the number of objects reclaimed (collected). The time between start and stop is the pause of that collection. If you place this next to the latency measured per request, you get the answer to "why did p99 jump?" When measured on this Pod with the fake service in the materials, collection ran about 300 times over 20,000 requests (around 2% of requests), and compared with the handler that breaks the cycles, p50 was similar while p99 jumped by several times. If more than 1% of requests take on a collection, that cost goes straight into p99. If you write the whole request time as the GC time, this causality disappears.
gc.freeze. According to the documentation, gc.freeze() moves all objects currently tracked into a permanent generation and ignores them in future collections. The use the documentation gives is before a fork: if you call gc.disable() early in the parent, gc.freeze() just before the fork, and gc.enable() early in the child, the child's collections do not touch the old objects inherited from the parent, so copying caused by copy-on-write is reduced. It also removes the cost of re-scanning, at every collection, the large static data built at startup.
Which line is growing. tracemalloc gives the location where a memory block was allocated and statistics per file and line, and lets you compute the difference between two snapshots to find leaks. Snapshot.compare_to(old, "lineno") returns results sorted by the absolute value of size_diff (the bytes that grew) in descending order. The top of statistics() for a single snapshot is "the line that takes up the most," so the big table built once at startup comes first. A leak is something that grows, so take a snapshot after warming up with a few requests, and take another after sending more, and look at the difference. The documentation says that the more frames you store, the greater the memory and CPU burden of tracemalloc itself.
The shallow size trap. sys.getsizeof counts only the memory directly attached to an object and does not count the objects it points to. The getsizeof of a single dictionary does not include the sizes of its keys and values, so to learn the size of a whole container, you have to follow references down, as in the recursive recipe linked from the official documentation, and avoid counting the same object twice. The slots section says that instances have a dictionary for attributes by default, which is wasteful for objects with only a few variables, so __slots__ can reduce that space. But when measured on this Pod, getsizeof actually shows the slots version as larger. That means what getsizeof measures differs from what is actually allocated, which is why the comparison is done by creating 100,000 objects with tracemalloc and dividing.
When a cache becomes a leak. A global dictionary filled with keys that never come back is simply a leak if it has no limit. Either put in a limit and discard the oldest first, or hold the value-side reference with a weakref so that it disappears when nothing else uses it. The weakref documentation says a weak reference alone cannot keep an object alive, and that caches holding large objects are the main use. It also says list and dict cannot be given weak references as they are, and int and tuple cannot even when subclassed, so you need to look first at what you store. A back-reference that creates a cycle, like a parent pointer, also loses the cycle if you change it to a weakref.
What it looks like in the field
On the JVM side, the "The heap had room, but the service stopped" course covers it with GC logs and heap dumps, and leak slopes from outside the process are covered by the "Diagnosing CPU and Memory Leaks" course. For an event loop server, a GC pause delays all requests together, just like a call that stalls the loop. The loop lag observation of the "It Wasn't One Request - Everything Got Slow" course applies here as it is.
What you will do in the next lab
You will print this Pod's GC thresholds, count the garbage left by a handler that creates cycles, and then break the cycles with a weakref. You will measure collection time with gc.callbacks and connect it to p99, and see how much gc.freeze reduces the full collection. After comparing getsizeof, deep size, and slots, you will find the leaking line with tracemalloc, fix it with a cache that has a limit, and confirm that the growth is gone.