TT Lab
Get started
Learn Learning paths Courses

One Slow Connection Froze Every Other One

Connection Probe: Concurrent Attempts and Safe Cancellation

Continue in TT Lab

Goal

Build a diagnostic tool that advances several TCP connection attempts with a limited budget, tells refusal, deadline, and cancellation apart, and reclaims all resources.

This is a 90-minute lab. The default session is 60 minutes, so extend it with +time before it expires (up to 180 minutes). Files disappear when the session ends. Keep important work separately before it ends.

Why it matters

A diagnostic tool that looks only at connection success hides the cause of failure, and a tool that cannot be canceled makes users force-kill it. This lab connects the earlier fairness, SO_ERROR, and control channel in a single loop. Real communication uses only the lab environment's loopback and does not inspect external servers. Data transfer, HTTP, TLS, retries, DNS, and the server's graceful shutdown are outside the scope of this lab. It is completion handling based on DefaultSelector and not an ET read server.

The checks import your source file and test its behavior. Do not start a server or wait for input at the top of the file. No additional packages or privileges are needed. The three probe.py files in /opt/fixtures/reactor are for separate conceptual observation and are not material files to modify or submit.

Steps

  1. Define the input and initial state of a connection attempt — In /root/reactor/client.py, implement Dial(sock, host, port, deadline=None). The input sock is a distinct Linux AF_INET/SOCK_STREAM socket opened by the caller. host accepts only a numeric IPv4 string, and port accepts only an int from 1–65535 excluding bool, and a violation is a ValueError. If deadline=None, it is time.monotonic()+30, and a specified value must be a finite, non-negative int/float excluding bool. Switch sock to non-blocking and keep sock, endpoint=(the normalized host, port), deadline as a float, state='new', and error=None. If creation fails, reclaiming the socket is the caller's responsibility.

  2. Tell an immediate success from the connecting state — In /root/reactor/client.py, Dial.start() calls sock.connect_ex(endpoint) once, only from new. 0/EISCONN is connected with error=0, EINPROGRESS/EALREADY is pending with error=None, and any other code is failed with error=that code. For an OSError, record the errno, and if there is no errno, it is EIO. If it is not new, return the existing state without calling. It returns the state string on every path, and in this Linux IPv4 contract, EAGAIN is not included in pending.

  3. Preserve the first SO_ERROR — In /root/reactor/client.py, Dial.writable() reads getsockopt(SOL_SOCKET, SO_ERROR) once only when pending, and records 0/EISCONN as connected with 0, and any other code as failed with that code. An OSError is handled as a failure with errno or EIO. If it is not pending, it does not query. It returns the current state. The caller must call this method for a pending Dial only after observing WRITE readiness. It also checks a real TCP listen port and a refused port that only did bind.

  4. Preserve the queue deadline and the cancellation reason — In /root/reactor/client.py, Dial.expire(now) changes to timed_out with ETIMEDOUT only when it is new/pending and now >= deadline. now is a finite monotonic time provided by the caller. Dial.cancel() changes only new/pending to cancelled with ECANCELED. Neither method overwrites a state and error that have already terminated, and neither closes the socket directly. They do not extend the deadline either.

  5. Build a ready queue without duplicates — In /root/reactor/client.py, implement push(item), pop(), discard(item), and len on ReadyQueue(). item is a hashable connection identifier that is not None. push ignores it if it is already there, pop removes and returns in FIFO order and returns None if empty, and discard removes that item and must be safe even if it is absent. A popped item can be added again, and identifiers of different generations of the same FD are distinguished. In run, use the Dial object itself as the key.

  6. Record the stop state and wake the kernel wait — In /root/reactor/client.py, Control() creates the reader/writer of a non-blocking socketpair and a threading.Event stop_event that is initially False. request_stop() calls writer.send(b'Q') once after Event.set and ignores only BlockingIOError. drain() calls reader.recv(4096) once and returns the bytes, but BlockingIOError gives b'' and EOF is a ConnectionError. close() closes the writer too even if an exception occurs while closing the reader. Calling request_stop after run has ended is outside the contract.

  7. Keep to the limit and budget and reclaim every socket — In /root/reactor/client.py, implement run(dials, control, limit=8, budget=4, ready=None). dials is a list of distinct new Dials that hold distinct open sockets (at most 128), limit is an int from 1–128 excluding bool, and budget is an int from 1–64 excluding bool. Validate the list type, the count, Dial duplicates, the new state, and the two budgets, and reject violations with a ValueError, in which case ownership remains with the caller. Once the DefaultSelector is successfully created, the function owns all the data sockets and the Control. Register the control reader with READ/data=None, and call ready once if it is given. On every turn, check for a stop first and apply the deadline to new/pending. If stopped, cancel all unfinished ones and finish. In one turn, start at most budget Dials and register at most limit pending Dials as WRITE/data=Dial. Reclaim items that terminated immediately. The wait is the smaller of the time to the nearest deadline and 1 second (minimum 0), and is 0 if the ready queue has items left or the waiting list has room to start more. Drain the control event from the select result and put WRITE events into the ready queue without duplicates. After checking for a stop again, take out at most budget items from the queue, check the deadline right before I/O, and call writable on the pending ones. Terminated items are deleted from the queue, unregistered, and then closed. Even if an exception occurs, it must reclaim the other data sockets, the Control, and the selector, and propagate the error. Do not let the failure of one reclamation block the rest. An empty list also reclaims the Control and returns []. A normal termination is a list in input order of {'state': , 'error': }. A diagnostic connection is closed even on success and is not reused.

  8. Aggregate successes and failure causes without loss — In /root/reactor/client.py, summarize(results) accepts only a list of dicts, and a violation is a ValueError. In each row, state must be connected/failed/timed_out/cancelled, and error must be an int excluding bool. connected is 0, the other states are positive, timed_out must be ETIMEDOUT, and cancelled must be ECANCELED. Unfinished ones such as pending are also rejected with a ValueError. The return is total, the count per each of the four states, and an errors dictionary. errors sums the errnos other than connected using the errno string as the key and sorts the keys. An empty input gives total and all four states as 0 and errors={}. It does not modify the original list or the rows. This step also rechecks the earlier real TCP, deadline, cancellation, budget, and exception reclamation.

Notes

Define the input and initial state of a connection attempt

Validate the address with ipaddress.IPv4Address without sneaking in name resolution. Note that bool is a subtype of int.

Tell an immediate success from the connecting state

Do not repeat reconnecting while you wait for write readiness. One Dial is one connection attempt.

Preserve the first SO_ERROR

SO_ERROR is cleared as it is read. Do not overwrite a failure you have already settled with the 0 of the next query.

Preserve the queue deadline and the cancellation reason

Time spent in the queue is also included in the total lifetime. Ending the state and actually reclaiming the FD are different steps.

Build a ready queue without duplicates

Manage the additions and removals of the deque and the set together. One job must not take another connection's turn with a duplicated place.

Record the stop state and wake the kernel wait

Record the stop state first and let notifications merge. Do not interpret one notification byte as one individual user command.

Keep to the limit and budget and reclaim every socket

Count the connection start budget and the completion handling budget separately. Both limits must hold in a test where completion events pile up all at once.

Aggregate successes and failure causes without loss

0 is an observation of success and not empty information. Tell the fact that aggregation is finished apart from the health of the whole production service.