One Slow Connection Froze Every Other One
An Operator's Connection Probe: Budgets, Causes, and Cleanup
In one line
Operable socket code has to be able to explain, with evidence, not just the bytes that succeeded but also the cause of termination, resource limits, and the reclamation on cancellation and exceptions.
Why this was needed
You ran the diagnostic tool ten times and got a result every time. But if two FDs are left behind on every run, a short demo succeeds while an operations tool left running for a long time eventually fails. Conversely, if you lump every error into a timeout, the graph looks plausible but you cannot tell a refused port from a user's cancellation. You have to check that no condition is missing between the test that lets the code pass and the evidence that helps the user's judgment.
This final module ties the connection state, ready queue, and control channel you learned earlier into one diagnostic tool. It is not a performance-contest tool or a scanner that sweeps arbitrary addresses. The class's communication is reproduced only on the loopback of your own lab environment and is not run against anyone else's server. You can create the important boundary conditions without adding external communication or capabilities.
How it works
Limits exist at several layers. run(limit=8) limits in-progress connections to at most 8. But if you opened the sockets in advance when creating the Dials, the remaining waiting entries hold FDs too. This lab limits the total list to 128 and counts the resources of the two control sockets and the selector separately. Since you wrote limit as 8, a report that the total FDs are 8 is wrong.
RLIMIT_NOFILE in Python resource is the limit on file descriptors a process can open. Raising it is not a fix for a leak. You first have to investigate by which path not only sockets but also files, control channels, and the selector are opened and closed. This class does not change host settings or process limits. The grader does not modify or delete the learner's files either.
Count the execution budgets separately
In one turn, start only budget new connections and handle only budget completion-ready jobs. The two are separate counters. If you test only an environment where every connection completes immediately, the two budgets look the same. The grader's deterministic scheduling test holds back completion events for the first two turns and then returns several at once. So it also exposes wrong code that limits starting but empties the whole completion queue.
The input list is finite, so even if the deadline check traverses the whole list, it stays within this limit. To scale this structure to hundreds of thousands of connections, you would also have to review a deadline heap or a timer structure and the memory cost of event batches. We do not recommend applying the small lab's linear traversal as is to a large-scale system.
Aggregate results exactly as they are
Each item leaves one of connected, failed, timed_out, or cancelled and an integer error. The 0 of connected is not an omission. Writing it like error or 기본값 (the placeholder is the default value) can turn a success into missing information. timed_out is this diagnostic tool's deadline overrun and cancelled is the user's interruption. The errno of failed is a clue for the operator to choose the next thing to investigate, but it alone does not settle whether the firewall, the application, or the network is the cause.
The final aggregation shows the total and the counts of the four states, and collects the number of cases per errno for the failure family separately. So that an unfinished pending is not quietly mixed into the denominator of the success rate, it is rejected at the aggregation input. This result is a summary of one run. It does not implement cumulative metrics over time or latency histograms.
| Evidence | What it confirms | What it does not confirm |
|---|---|---|
| Comparing real TCP listen/bind | Connection establishment success and refusal | TLS, HTTP, business processing |
| Waiting on a full socketpair | Deadline and control wake-up | Real remote SYN loss |
| Synthetic batches of completion events | Per-turn budget and ready queue order | The kernel's fairness guarantee |
| Injecting a close error | Whether it cleans up the later resources too | Every combination of OS errors |
The grader's PendingSocket is not a remote server with a genuinely delayed connection. It only returns an in-progress result for connect, and creates the condition of having no write readiness with a full local socket. The reason for leaving this distinction in the text is to avoid exaggerating a mock test result as real network validation.
A single finally is not enough
If an exception occurs on the first close of the cleanup loop, the later sockets may not be closed. You have to try reclaiming each socket, keep the first cleanup error, and then clean up the rest and even the control channel. Do not hide the error; propagate it at the end. Check not only whether a normal result was returned but also whether any open FD remains after a failure. When removing from the registration list, too, unregister comes before close.
What it looks like in the field
In incident reports, write the conditions instead of "tests passed." Record the number of targets, the concurrency limit, the work budget, the counts of successes, refusals, cancellations, and deadlines, and whether FDs were reclaimed after it ended. You cannot call something fast just because CPU is low, and you cannot say small requests were handled fairly just because throughput is high. This lab is a starting point for designing operational metrics and does not guarantee a real service's SLO or load capacity.
What you will do in the next lab
In the 90-minute combined lab, you complete the connection diagnostic tool in 8 steps. You look together at whether the correct answer passes and whether the wrong answers, each with one contract broken, fail. Files disappear when the session ends, so extend it with +time and keep important work separately before it ends.