TT Lab
Get started
Learn Learning paths Courses

Debugging in Practice

It dies every afternoon — catching an exhausted resource with data

Continue in TT Lab

Goal

You prove a file descriptor leak with a slope, lower the limit to bring the floor forward and reproduce it, and record the errors that arise in unrelated places after hitting the floor. You also reproduce memory safely with an address-space limit and write down the difference between what is received as an exception and what gets killed, and prove with the same tool that the fixed version's slope is 0.

Why it matters

Resource exhaustion is a typical incident where the symptom does not point to the cause. It was the session handler that leaked the descriptors, yet the error comes from the database connection. That is because the code that leaked has already taken its share, and the moment the floor is exposed, the code that requested a resource next fails. So this investigation is not about reading messages but about measuring. If you vary the number of requests and measure the number of open descriptors, you get a few points, and their slope is the number leaked per request. A slope of 0 means no leak. This is how you turn a "feeling" into data. Limits are also an investigation tool. A process can lower its soft ceiling on its own, so you can create in a few seconds right now the floor you would hit ten hours from now. Conversely, if you cover it up by raising the limit, only the time of death is pushed back. The grader does not trust your conclusions. It builds a separate handler that knows exactly how many it leaks per request, actually hooks your instrumentation tool up to it, and directly compares the slope, the limit, and the error name. That count changes on every run.

Steps

  1. Create and run /root/exhaust/gen_exhaust.py to create /root/exhaust/leaky.py.
  2. Use /root/exhaust/fdcount.py to count the open descriptors of a live process.
  3. Use /root/exhaust/measure_leak.py to measure while varying the number of requests and leave the slope in /root/exhaust/trend.json.
  4. Use /root/exhaust/run_under_limit.py to lower the limit and bring the floor forward, and write it in /root/exhaust/nofile.json.
  5. Use /root/exhaust/symptoms.py to collect the failures after hitting the floor and write them in /root/exhaust/symptoms.json.
  6. Use /root/exhaust/mem.py to reproduce the address-space limit and write it in /root/exhaust/mem.json.
  7. Measure the fixed version with the same tool and leave a slope of 0 and a pass in /root/exhaust/fixed.json.
  8. Report in /root/exhaust/summary.json and /root/exhaust/exhaust_report.md in four sections.

Notes

Get the session handler in hand

Create and run /root/exhaust/gen_exhaust.py to create /root/exhaust/leaky.py. You give the number of requests with --requests and can run the fixed version with --fixed.

Just save this script as it is and run it. Run leaky.py once with --requests 100. Nothing seems to happen — descriptors pile up only inside the process, and when the process ends, the kernel reclaims them all.

Count how many it is holding right now

Create /root/exhaust/fdcount.py so that, given --pid <번호> (the process ID), it counts that process's open descriptors and outputs JSON containing pid and open_fds.

Linux places a /proc//fd directory for each process in proc(5) and puts one entry per open descriptor. Counting is counting the entries in that directory. Also decide in advance how to answer when asked about a process that does not exist.

Work out the slope: how many leak per request

Use /root/exhaust/measure_leak.py to measure the descriptor count at 10, 40, and 80 requests, and leave target, points, per_request, and baseline in /root/exhaust/trend.json. per_request must be at least 1.

The handler creates the ready file and then stays alive until the pause file appears. You can count via /proc in between. Two points give a slope, but with three you can also see whether it is a line. Do not forget to create the pause file after measuring to let the handler go.

Lower the limit to bring the floor forward

Use /root/exhaust/run_under_limit.py to run leaky.py under a low --nofile and leave nofile, cmd, exit_code, errno_name, stderr_tail, and stdout_tail in /root/exhaust/nofile.json. errno_name must be EMFILE.

resource.setrlimit applies to its own process and the children created afterward. If you lower it in subprocess's preexec_fn, you can put only that command in a narrow room. Give a number of requests comfortably larger than the limit — you only see the floor if you reach the limit.

The symptom does not point to the cause

Use /root/exhaust/symptoms.py to exhaust the descriptors, then try four things, open, socket, subprocess, and sqlite3, and leave nofile, held, observations, and misleading_ops in /root/exhaust/symptoms.json. misleading_ops are the items whose error message says nothing about descriptors.

All four need a descriptor, but the faces of the messages differ. In particular, sqlite3 says only 'unable to open the database file.' If you see that message and dig through disks and permissions, hours vanish. Leave a list of which items hide the cause.

How does memory run out

Use /root/exhaust/mem.py to reproduce an allocation that exceeds the address-space limit and one that does not, and leave the two results, over and under, and a note in /root/exhaust/mem.json. The outcome of over must be MemoryError, and under must be ok.

RLIMIT_AS is a ceiling on the address space, so an allocation that exceeds the limit comes up as MemoryError. It matters that a traceback is left — the kernel's OOM killer uses SIGKILL and leaves no trace at all. In the note, write this difference and the fact that address space differs from actual usage.

Prove that the fixed version's slope is 0

Measure the fixed version (--fixed) with the same tool and leave per_request_before, per_request_after, nofile, requests, and exit_code_after in /root/exhaust/fixed.json. per_request_after must be 0 and exit_code_after must be 0.

The claim that you fixed it must also be made with the same instrumentation. A slope of 0 means descriptors do not grow even as requests increase, and if it processes many requests under a low limit and still passes, it means it never hits the floor. Leave both pieces of evidence together.

Report the symptoms and the cause separately

In /root/exhaust/summary.json, write per_request_before, per_request_after, nofile, errno_name, misleading_ops, and mem_outcome, and in /root/exhaust/exhaust_report.md, report in four sections: ## 무엇이 바닥났나 ## 증상은 어디서 났나 ## 어떻게 증명했나 ## 남은 위험 (in order, these mean: what ran out, where the symptoms arose, how it was proved, and the remaining risk).

The value of the report lies not in 'descriptors leaked' but in 'one leaked per request, and after the fix it is 0.' Write the list of symptoms and the cause separately so that the next person does not dig through disks after seeing the sqlite message.