CS for Building Good Services — Relearning Textbook Ideas by Measuring
Find GC Pauses and the Leaking Line with Numbers
Goal
You will print this Pod's GC settings, count the garbage left by a handler that creates reference cycles, and then break the cycles with a weakref. You will measure GC pauses with gc.callbacks and connect them to the p99 of request latency, and see the effect of gc.freeze. You will confirm the trap of sys.getsizeof and slots with tracemalloc, find the line (file:line) that grew in a leaking service, fix it with a cache that has a limit, and measure whether the growth is gone.
Why it matters
CPython clears most objects immediately with reference counting, but cycles are cleared in batches by the generational collector according to the number of allocations. If that collection lands on more than 1% of requests, the average stays the same but p99 jumps. A memory problem has to be looked for not at "the biggest line" but at "the line that grew the most," and getsizeof does not count the objects it points to, so it reports sizes that are too small. The grader does not look only at the numbers you wrote; it runs your functions again in a fresh interpreter on inputs that are not the materials and compares them with the reference.
Materials
They are under /opt/fixtures/svccs/memory/. Only read them. From Python, call it with sys.path.insert(0, "/opt/fixtures/svccs/memory") followed by import svc.
svc.py 가짜 주문 서비스
CATALOG 불러올 때 한 번 만드는 상품 2만 개(자라지 않는 큰 정적 데이터)
load_requests() requests.jsonl 의 요청 목록
handle_cyclic 요청마다 부모↔자식이 서로 가리키는 트리를 만든다 → {"id", "items", "path"}
render(req) 요청을 문자열로 렌더링
handle_leaky 렌더링 결과를 전역 사전에 넣고 지우지 않는다 → {"id", "bytes"}
requests.jsonl 요청 3,000줄 {"id", "user", "items"} — id 는 다시 오지 않는다
sessions.json 세션 400개 {세션 id: {"user", "roles", "cart", "flags"}}
Steps
- In
/root/svccs/memory/gcinfo.json, writepython(platform.python_version()),threshold(gc.get_threshold()as a list), andgc_enabled(gc.isenabled()). - Run the first 200 of
svc.load_requests()in the ordergc.collect()→gc.disable()→svc.handle_cyclic200 times →gc.collect()(the returned value is the number of unreachable objects reclaimed) →gc.enable(), and in/root/svccs/memory/cycles.jsonwriterequests·unreachable·per_request(unreachable ÷ requests, to two decimal places). - In
/root/svccs/memory/mem.py, createhandle_acyclic(req). It returns the same value assvc.handle_cyclic(req)but does not create cycles (child → parent through aweakref). The grader calls it many times with GC turned off and checks whethergc.collect()returns 0. - In the same file, create
measure(handler, reqs)following the "Measurement rule" below. Extend the request list to[base[i % len(base)] for i in range(20000)](base =svc.load_requests()), and in/root/svccs/memory/pause.jsonwrite{"cyclic": measure(svc.handle_cyclic, …), "acyclic": measure(mem.handle_acyclic, …)}. - In a process that has imported
svc, measure the shortest of three runs ofgc.collect()(in ms), then, right aftergc.freeze(), thegc.get_freeze_count(), and again the shortest of three runs ofgc.collect(), and thengc.unfreeze(). In/root/svccs/memory/freeze.json, write them ascollect_ms_before·freeze_count·collect_ms_after. - In the same file, create
deep_size(obj)(the sum ofsys.getsizeofover each distinct object (by id), following the keys and values of a dict and the elements of list, tuple, set, and frozenset),PlainPoint(a, b, c, d)with four attributes a·b·c·d, a class with the same attributes held in__slots__,SlotPoint(a, b, c, d), andper_object_bytes(cls, n=100000)(turn tracemalloc on, create n instances ofcls(0, 0, 0, 0)into a list, and take the difference ofget_traced_memory()[0]before and after ÷ n, to one decimal place). Load it withjson.loadfromsessions.json, and in/root/svccs/memory/sizes.jsonwriteshallow(getsizeof)·deep·plain_getsizeof·slots_getsizeof(the getsizeof of a single instance of(0, 0, 0, 0)for each)·plain_per_object·slots_per_object. - Before importing
svc, calltracemalloc.start(), send the first 500 requests throughhandle_leakyand take a snapshot, then send the next 2,000 (the 500th to 2,499th) and take a snapshot. Take the first entry ofcompare_to(앞 스냅숏, "lineno")(the previous snapshot) and put it in/root/svccs/memory/leak.jsonasfile(file name only)·line·size_diff_kb(size_diff ÷ 1024, to one decimal place)·count_diff. - In the same file, create
handle_fixed(req). It returns the same value assvc.handle_leakybut limits the cache to at most 1,000 entries (discarding the oldest first when it overflows). Turn tracemalloc on, warm up each handler with the first 1,500 requests, and then measure how muchget_traced_memory()[0]grew (in KB, to one decimal place) over the next 1,500 requests. In/root/svccs/memory/fix.json, writerequests_measured(1500)·leaky_growth_kb·fixed_growth_kb·leak_line(the "file:line" of step 7).
Measurement rule
시작 전에 gc.collect() 한 번. gc.callbacks 에 콜백을 넣어 phase "start" 에서 시각을 적고
"stop" 에서 그 차이를 일시정지 하나로 기록한다. 요청마다 handler(req) 앞뒤를 perf_counter 로 잰다.
끝나면 콜백을 뺀다(예외가 나도 — try/finally).
돌려줄 것: requests · collections(일시정지 개수) · gc_pause_ms_total · gc_pause_ms_max(소수 셋째 자리)
total_ms(요청 지연의 합, 소수 셋째 자리)
p50_ms · p99_ms(지연을 정렬한 목록의 [int(n×0.5)] · [int(n×0.99)] 번째, 소수 넷째 자리)
Notes
- The grader imports
mem.py. Put code that produces the result JSON underif __name__ == "__main__":or in a separate script. - GC and tracemalloc numbers depend on what that process has already imported. Measure in a fresh
python3process for each step, in the order given. - Common mistakes: writing the size of a whole container from getsizeof alone, naming the biggest line of a single snapshot as the culprit, writing the whole request time as the GC pause, and setting the limit too large so that growth remains even after the fix.
- The outputs disappear when the session ends. Keep them elsewhere if you need them.
Print this Pod's GC settings
Print python, threshold, and gc_enabled with this Pod's python3 and write them to /root/svccs/memory/gcinfo.json.
gc.get_threshold() returns a tuple, so write it as a list. Threshold values differ between versions, so the value printed on this interpreter is the reference, not the number in the documentation.
Count the garbage that reference counting does not clear
Call svc.handle_cyclic 200 times with GC turned off, and write the number reclaimed by gc.collect() to /root/svccs/memory/cycles.json as requests, unreachable, and per_request.
An object without a cycle disappears the moment the request ends because its reference count reaches 0, so with GC turned off, what remains is only the cycles. Clear the earlier garbage first with gc.collect() before starting, so that you count only what these requests left behind. If you run with GC on, an automatic collection takes some of it in the middle.
Break the cycles with weakref
Create handle_acyclic(req) in /root/svccs/memory/mem.py: the same value as svc.handle_cyclic, no cycles. The grader compares values using other requests, and checks whether gc.collect() returns 0 after running with GC turned off.
Keep parent → child (the children list) as a strong reference, and hold only child → parent with weakref.ref(parent), and the cycle disappears. When you walk back up the path, call the weak reference (ref()) to get the parent. The root is held by a local variable until the function ends, so it does not disappear in the middle.
GC pauses make p99
Create measure(handler, reqs) in mem.py following the measurement rule, measure the two handlers with 20,000 requests, and write cyclic and acyclic to /root/svccs/memory/pause.json. The grader calls your measure again with its own handlers (one with no collection, one that calls gc.collect).
A pause is the time between the callback's start and stop, not the whole request time. Collection is triggered by the number of allocations, so only the side that creates cycles runs it often. If collection lands on more than 1% of requests, p99 takes on that cost.
Set aside long-lived objects with gc.freeze
In a process that has imported svc, measure the full collection time and freeze_count before and after freeze, and write them to /root/svccs/memory/freeze.json as collect_ms_before, freeze_count, and collect_ms_after.
freeze moves the objects being tracked "right now" into a permanent generation. So you must call it after building the large static data (CATALOG in svc) for it to have an effect. Times wobble, so use the shortest of three runs.
getsizeof is shallow: deep size and slots
Create deep_size, PlainPoint, SlotPoint, and per_object_bytes in mem.py and write /root/svccs/memory/sizes.json. The grader compares deep_size on other objects that mix sharing and cycles.
getsizeof does not count the keys and values of a dictionary. You follow references down, but you must skip objects already counted by id so that shared strings are not counted twice and it terminates even on a list that contains itself. A slots instance can look larger by getsizeof, so trust the value from creating 100,000 objects with tracemalloc.
Not the biggest line, but the line that grew the most
Turn tracemalloc on before importing svc, and from the snapshot difference between after warming up with 500 handle_leaky requests and after sending 2,000 more, write the first line to /root/svccs/memory/leak.json as file, line, size_diff_kb, and count_diff.
The top of statistics for a single snapshot is the line that "takes up the most," which gives the large data built once at startup. A leak is something that "grows" over time, so look at the difference of two snapshots with compare_to. filename is a full path, so keep only the name with os.path.basename.
Fix it with a limited cache and measure whether the growth is gone
Create handle_fixed(req) in mem.py with the cache limited to at most 1,000 entries, measure the growth of the two handlers, and write requests_measured, leaky_growth_kb, fixed_growth_kb, and leak_line to /root/svccs/memory/fix.json. The grader re-measures your handle_fixed in a new process with new request ids.
A cache filled with keys that never come back is a leak if it has no limit. Put it in an OrderedDict and, when the length exceeds the limit, discard the oldest with popitem(last=False). If the number of warm-up requests is smaller than the limit, the cache is still growing during the measurement interval, so growth appears to remain. The grader warms up with 2,000 requests using new request ids and then measures 4,000.