It dies every afternoon — catching an exhausted resource with data
Goal
You prove a file descriptor leak with a slope, lower the limit to bring the floor forward and reproduce it, and record the errors that arise in unrelated places after hitting the floor. You also reproduce memory safely with an address-space limit and write down the difference between what is received as an exception and what gets killed, and prove with the same tool that the fixed version's slope is 0.
Why it matters
Resource exhaustion is a typical incident where the symptom does not point to the cause. It was the session handler that leaked the descriptors, yet the error comes from the database connection. That is because the code that leaked has already taken its share, and the moment the floor is exposed, the code that requested a resource next fails. So this investigation is not about reading messages but about measuring. If you vary the number of requests and measure the number of open descriptors, you get a few points, and their slope is the number leaked per request. A slope of 0 means no leak. This is how you turn a "feeling" into data. Limits are also an investigation tool. A process can lower its soft ceiling on its own, so you can create in a few seconds right now the floor you would hit ten hours from now. Conversely, if you cover it up by raising the limit, only the time of death is pushed back. The grader does not trust your conclusions. It builds a separate handler that knows exactly how many it leaks per request, actually hooks your instrumentation tool up to it, and directly compares the slope, the limit, and the error name. That count changes on every run.
Steps
- Create and run /root/exhaust/gen_exhaust.py to create /root/exhaust/leaky.py.
- Use /root/exhaust/fdcount.py to count the open descriptors of a live process.
- Use /root/exhaust/measure_leak.py to measure while varying the number of requests and leave the slope in /root/exhaust/trend.json.
- Use /root/exhaust/run_under_limit.py to lower the limit and bring the floor forward, and write it in /root/exhaust/nofile.json.
- Use /root/exhaust/symptoms.py to collect the failures after hitting the floor and write them in /root/exhaust/symptoms.json.
- Use /root/exhaust/mem.py to reproduce the address-space limit and write it in /root/exhaust/mem.json.
- Measure the fixed version with the same tool and leave a slope of 0 and a pass in /root/exhaust/fixed.json.
- Report in /root/exhaust/summary.json and /root/exhaust/exhaust_report.md in four sections.
Notes
- Handler contract:
python3 /root/exhaust/leaky.py --requests N [--fixed] [--dir D] [--ready-file F] [--pause-file P]processes the requests, creates the ready file, stays alive until the pause file appears, and then prints one line of JSON. While it is alive, you can measure the descriptor count. - Counting tool:
python3 /root/exhaust/fdcount.py --pid <번호>(where the placeholder is the process ID) outputs one JSON object containing pid and open_fds. Linux places one entry per open descriptor in/proc/<pid>/fd. - Instrumentation tool:
python3 /root/exhaust/measure_leak.py --target <처리기> --points 10,40,80 [--target-arg=--fixed] --out <json>(the placeholder is the handler path) outputs target, target_args, points (pairs of request count and descriptor count), per_request (the slope), and baseline. Values given with--target-argare passed to the handler as they are (write them joined with an equals sign). The slope is (last descriptor count - first descriptor count) / (last request count - first request count). - Limit tool:
python3 /root/exhaust/run_under_limit.py --nofile <상한> --cmd "<명령>" --out <json>(the placeholders are the ceiling and the command) runs only that command under a lowered ceiling and outputs nofile, cmd, exit_code, errno_name, stderr_tail, and stdout_tail. errno_name isEMFILEif descriptors ran out, otherwise null. - Symptom tool:
python3 /root/exhaust/symptoms.py --nofile <상한> --out <json>(with the ceiling filled in) exhausts the descriptors, then tries four things, open, socket, subprocess, and sqlite3, and outputs nofile, held, and observations (op, error, errno). - Memory tool:
python3 /root/exhaust/mem.py --as-mb <상한> --alloc-mb <할당> --out <json>(with the ceiling and the allocation size filled in) outputs as_mb, alloc_mb, outcome (MemoryError or ok), and exit_code. resource.setrlimitapplies to its own process and the children created afterward. If you lower it right before launching the child (preexec_fn), you can put only that command in a narrow room while leaving the parent as it is.- Common mistakes: measuring only once and calling it a leak, covering it up by raising the limit, counting descriptors as files only, and trying to read /proc of a process that has already ended.
- Assumption of this lab: treating the slope as a straight line is because this handler has a simple structure that opens the same number for every request. In a real service, you must add more points and first check whether it is a line.
- Do not build a load test. The Pod has 2 cores and 2Gi of memory, and the budget for one grading is 60 seconds.
Get the session handler in hand
Create and run /root/exhaust/gen_exhaust.py to create /root/exhaust/leaky.py. You give the number of requests with --requests and can run the fixed version with --fixed.
Just save this script as it is and run it. Run leaky.py once with --requests 100. Nothing seems to happen — descriptors pile up only inside the process, and when the process ends, the kernel reclaims them all.
Count how many it is holding right now
Create /root/exhaust/fdcount.py so that, given --pid <번호> (the process ID), it counts that process's open descriptors and outputs JSON containing pid and open_fds.
Linux places a /proc//fd directory for each process in proc(5) and puts one entry per open descriptor. Counting is counting the entries in that directory. Also decide in advance how to answer when asked about a process that does not exist.
Work out the slope: how many leak per request
Use /root/exhaust/measure_leak.py to measure the descriptor count at 10, 40, and 80 requests, and leave target, points, per_request, and baseline in /root/exhaust/trend.json. per_request must be at least 1.
The handler creates the ready file and then stays alive until the pause file appears. You can count via /proc in between. Two points give a slope, but with three you can also see whether it is a line. Do not forget to create the pause file after measuring to let the handler go.
Lower the limit to bring the floor forward
Use /root/exhaust/run_under_limit.py to run leaky.py under a low --nofile and leave nofile, cmd, exit_code, errno_name, stderr_tail, and stdout_tail in /root/exhaust/nofile.json. errno_name must be EMFILE.
resource.setrlimit applies to its own process and the children created afterward. If you lower it in subprocess's preexec_fn, you can put only that command in a narrow room. Give a number of requests comfortably larger than the limit — you only see the floor if you reach the limit.
The symptom does not point to the cause
Use /root/exhaust/symptoms.py to exhaust the descriptors, then try four things, open, socket, subprocess, and sqlite3, and leave nofile, held, observations, and misleading_ops in /root/exhaust/symptoms.json. misleading_ops are the items whose error message says nothing about descriptors.
All four need a descriptor, but the faces of the messages differ. In particular, sqlite3 says only 'unable to open the database file.' If you see that message and dig through disks and permissions, hours vanish. Leave a list of which items hide the cause.
How does memory run out
Use /root/exhaust/mem.py to reproduce an allocation that exceeds the address-space limit and one that does not, and leave the two results, over and under, and a note in /root/exhaust/mem.json. The outcome of over must be MemoryError, and under must be ok.
RLIMIT_AS is a ceiling on the address space, so an allocation that exceeds the limit comes up as MemoryError. It matters that a traceback is left — the kernel's OOM killer uses SIGKILL and leaves no trace at all. In the note, write this difference and the fact that address space differs from actual usage.
Prove that the fixed version's slope is 0
Measure the fixed version (--fixed) with the same tool and leave per_request_before, per_request_after, nofile, requests, and exit_code_after in /root/exhaust/fixed.json. per_request_after must be 0 and exit_code_after must be 0.
The claim that you fixed it must also be made with the same instrumentation. A slope of 0 means descriptors do not grow even as requests increase, and if it processes many requests under a low limit and still passes, it means it never hits the floor. Leave both pieces of evidence together.
Report the symptoms and the cause separately
In /root/exhaust/summary.json, write per_request_before, per_request_after, nofile, errno_name, misleading_ops, and mem_outcome, and in /root/exhaust/exhaust_report.md, report in four sections: ## 무엇이 바닥났나 ## 증상은 어디서 났나 ## 어떻게 증명했나 ## 남은 위험 (in order, these mean: what ran out, where the symptoms arose, how it was proved, and the remaining risk).
The value of the report lies not in 'descriptors leaked' but in 'one leaked per request, and after the fix it is 0.' Write the list of symptoms and the cause separately so that the next person does not dig through disks after seeing the sqlite message.