One Slow Connection Froze Every Other One
When do you drop a silent connection
In one line
A quiet connection and a dead connection cannot be told apart by looking at the socket alone, so the program has to write down the last activity time itself.
Why this was needed
A multiplexing server is built to hold connections for a long time. That is both its strength and its problem. If a client sends a request, gets an answer, and then says nothing, the server keeps watching that connection. One file descriptor, the receive and send buffers, and the per-connection state the program created all stay in place. When such connections pile up, at some point you can no longer accept new ones, and the error you see then is usually resource exhaustion, which points to a cause far from the real one.
The worse case is when the other side has already vanished. If a laptop lid is closed or a device in the middle forgets the state, the other side cannot send a FIN or an RST. The server-side socket stays open, and read readiness never arrives. This state is called half-open. Querying the socket state alone cannot tell you whether it is quiet or dead.
How it works
The solution is to record time instead of asking the socket. For each connection, write down when something meaningful last happened, compare it with the current time periodically, and cut off the old ones.
conn.last_active = clock() # 읽거나 보낸 뒤에 갱신
stale = [key for key, conn in conns.items()
if now - conn.last_active >= idle_timeout]
You have to decide three things: what counts as activity, how often to check, and which clock to use.
The definition of activity differs by service. You can count only receiving bytes as activity, or include sending a response too. What matters is to gather the definition in one place. If several places update the time in their own ways, you will not be able to explain later why a connection was not cut.
The check interval is tied to the wait time of the multiplexing call. If there are no events at all, the program is asleep inside select, so however accurate the expiry list you build, nobody gets cut off unless it wakes up. So you put an upper bound on the wait time. It makes the program wake periodically even without events and sweep once around.
The clock is for measuring elapsed time, so you use a monotonic clock, not a wall clock. If time synchronization adjusts the wall clock backward, a connection that has not yet expired can look expired, or conversely may never expire. If you make the clock injectable for testing, you can test the expiry decision without actually waiting several seconds.
The kernel has a similar mechanism. The TCP keepalive described in tcp(7) and socket(7) sends an empty probe on a connection that has been quiet for a long time to check that the other side is alive. But this is a liveness check at the transport layer, not an application's idle policy. Even if the other side's process is stopped, the kernel can still answer, so keepalive passes and yet no work proceeds. The two mechanisms do not substitute for each other.
There is an order for cutting a connection. Remove it from the interest list, erase the per-connection state the program held, and close the socket. If you close only the socket and leave it in the list, the next wake-up looks for a connection that does not exist, and conversely if you remove it only from the list without closing it, the descriptor leaks. In the lab, you gather this cleanup into one function.
Separating the choosing of expiry targets from the actual cutting is for the same reason. If you delete while choosing, the data structure changes in the middle of the traversal and some connections get skipped. A more practical reason is testing. If the choosing function is pure, you can check boundary conditions just by feeding in a clock, without opening real sockets or waiting several seconds. Questions like whether to cut when the elapsed time is exactly equal to the limit can only be pinned down this way.
When you leave a log after cleanup, write the reason too. A connection cut because the other side closed it and one we cut because of the idle limit have completely different causes, and if the log says only "connection closed," there is no way to tell the two apart later. This distinction is exactly the basis for deciding in production whether to adjust the idle limit.
What it looks like in the field
A combination often seen in production looks like this. The connection count graph rises monotonically without sawteeth, and requests per second are the same as usual. That is, new clients keep arriving and old clients do not leave. This kind of graph shows up when there is no idle cleanup, or when there is one but it does not wake up when there are no events.
An incident in the opposite direction from setting the idle limit too short is also common. Clients that were built to reuse connections end up opening a new connection every time, so latency grows and ports are used up quickly. So the number has to be decided together with the clients' reuse interval. An idle limit chosen by looking only at the server is usually wrong on one side.
What you will do in the next quiz
Sort out what tells a quiet connection from a cut connection, how the expiry check runs when there are no events at all, and what keepalive guarantees and what it does not. The idle_keys you will build in the lab only chooses and does not delete, and it is good to think about that reason too.