TT Lab
Get started
Learn Learning paths Courses

One Slow Connection Froze Every Other One

The Cancel Button Does Nothing: Control Channels and Ownership

Continue in TT Lab

In one line

Recording the intent to stop and waking a loop that is asleep in the kernel are separate things. Record the state first, notify through the control socket, and have the loop that owns the sockets do the cleanup.

Why this was needed

You put a cancel button on a diagnostic screen. Another thread wrote stop=True, but the loop is waiting inside select. The user feels the button doesn't work and presses it several times. Even though the value in memory has already changed, if the loop cannot execute its next line, there is no chance to see that value. If you change the select wait to be very short, the response gets faster, but it has to wake up repeatedly even when idle.

What you need then is the two parts of "state and notification." The state preserves the fact that cancellation was requested, and the notification creates a turn to read that fact. The mere fact of receiving a notification does not let you judge that the cancellation has been handled or that all resources have been reclaimed. If you think of request, observation, and cleanup completion separately, the boundary between the UI and the backend becomes clear too.

How it works

threading.Event holds a flag and notifies threads waiting with Event.wait. It is not a device that automatically puts an event into a selector waiting for socket readiness. In the lab, you register the reading end of a socketpair with the selector and leave the writing end with the requester. You can watch the completion of data connections and the control notification together in a single wait.

python3 /opt/fixtures/reactor/wakeup_probe.py

The experiment prepares a thread that goes into select(2) and sets the Event. During the short observation, the thread does not end. Then, when it sends one byte, Q, to the control socket, read readiness appears, and the loop wakes up and checks the Event that is already set. This observation time is not a guaranteed latency in a production environment but a condition for creating the counterexample that a state change alone does not wake a socket wait.

State must come first

If you send the notification first and set the Event later, the loop may wake up in between. After it sees that it is not yet a stop and goes back to sleep, there may be no new notification. So request_stop goes in the order Event.set and then writer.send. Make both ends of the control socket non-blocking. If a cancel request on the UI side stops until the send buffer empties, the control channel becomes a new bottleneck.

If the buffer is full and BlockingIOError occurs, it means there is already data to read. The stop state remains in the Event, so you do not retry the same notification endlessly. It is a design where several requests may merge into a single wake-up. It differs from a command queue that maps one byte to one user request. For a channel where every message has to be processed, like order or cancellation IDs, you need a separate queue and framing.

drain reads at most 4096 bytes just once. If you read without limit to empty out the notifications, the thread pouring in notifications can starve data processing. The remaining notifications are handled on the next turn. EAGAIN for an empty state and EOF for channel termination are different too. EAGAIN means there is nothing to read right now, and EOF means the other side of the control connection has closed. The lab handles an unexpected EOF as a ConnectionError and reclaims all resources.

Tell the notification count from processing completion

Even if the user pressed cancel three times, one termination handling is enough. When you show "cancellation complete" on the response screen, you should check whether the run has ended, not the number of bytes sent. Conversely, if a thread waiting for the completion notice closes the control writer first, EOF occurs and it may enter an exception path different from a normal cancellation. If you draw a short timeline of the order in which ownership is passed among the terminate button, the request thread, and the loop, it is easy to find this kind of race.

Who closes

The request thread does not arbitrarily close the data sockets. The loop has to manage those sockets' registration state and the ready queue together. The cancel requester uses only the Event and the control writer, and the loop changes unfinished Dials to cancelled and then does unregister and close. It follows the unregister-then-close order of the selectors documentation.

In this API, when you call run, it takes over the sockets and the Control of the valid input, and you do not use them again after it returns or raises. If invalid arguments are rejected before running, ownership still belongs to the caller. Calling request_stop again after run has ended is not allowed. When you connect it to a real screen, you need a lifecycle agreement that shares the completion state and does not release the channel before the requester finishes.

What it looks like in the field

If a production tool does not stop even after it receives a termination request, users end up force-killing the process. A server that needs to preserve data needs more: rejecting new requests, draining in-flight requests, and forced termination after a time limit. This one is the immediate cancellation of a connection diagnostic tool. You must not describe it as having implemented a server's full graceful shutdown.

What you will do in the next check

In the next quiz, you judge the order of state and notification. After that, in the combined lab, you build the Control and then wake, from another thread, a loop that has entered a real kernel wait. Tell apart the wrong answer that only changes the stop flag from the wrong answer that leaves the control socket blocking. The final judgment is not a log saying the button was pressed but the cancelled classification of unfinished results and the reclamation of every FD.