TT Lab
Get started
Learn Learning paths Courses

One Slow Connection Froze Every Other One

One person stops the whole line

Continue in TT Lab

In one line

In a sequential server, if one connection starts talking late, every connection lined up behind it waits that same time.

Why this was needed

Before talking about thousands of connections, you have to see how two connections block each other. The most common server shape is a loop that takes one connection with accept, handles it to the end, and goes back to accept. This structure is easy to read, and it has no problem when clients arrive one at a time. The problem shows up not when clients overlap, but when the client in front is slow.

"Slow" here splits into two kinds. One is that the server takes a long time to process the request, and the other is that the client starts talking late. The second is more dangerous. The server is doing nothing, yet the whole process is stopped inside recv, waiting for the other side to send its first byte. At this point CPU usage is close to zero, and no line is left in the log. Busy and blocked look almost the same on the surface metrics.

If the previous course covered byte boundaries and reconnection, this one covers a question that comes before that. What has to change for a single process to hold many connections at the same time?

How it works

The socket calls that can stop are fixed. In blocking mode, all three below stop the process.

Call When it stops Meanwhile, other connections
accept When the listen queue is empty Cannot accept new connections
recv When there are no received bytes Cannot read even connections already accepted
send When there is no room to send Cannot send responses

New connections do not disappear entirely. The kernel keeps connections that have finished the 3-way handshake in the listen queue and hands them over one at a time when accept is called. So from the client's side, connect appears to succeed right away, and only the response never comes. On screen, this shows up as "it connects but it's slow."

The size of that queue is the backlog of listen, and listen(2) says that if this value is larger than /proc/sys/net/core/somaxconn, it is silently capped to that value. According to the same document, the default of that file is 4096 from Linux 5.4 and 128 on earlier kernels. That is, even if you write a big number in code, the kernel setting decides the actual queue. And enlarging the queue only increases the number of people waiting; it does not shorten the waiting time.

Measuring comes first. If you attach one slow client first and line up six fast clients behind it, in a sequential server all six wait as long as the slow client takes. The measuring tool used in the lab reports that number like this.

client 1: 1699 ms OK
...
blocked=6
fast=0

blocked is the number of clients that took more than 1 second, and fast is the number of clients that finished within 0.2 seconds. When you switch to multiplexing, these two numbers flip under the same load. If you don't measure before the change, you cannot say what improved, and you cannot tell when nothing improved.

Blocking itself is not a bad design. The Python Socket HOWTO also starts its explanation with blocking sockets. However, in blocking mode you can handle only one connection at a time, and to handle several connections at once you either give each connection its own flow of execution (threads or processes) or switch to calls that do not stop. This course takes the second road.

Giving each connection a thread is also widely used in practice. It has the big advantage that the code keeps its blocking shape, so it is easy to read. However, each thread has its own stack and context switches cost something, so if the goal is thousands of connections, you reach the limit on the resource side before the throughput side. More important is that the two roads solve the same problem differently. Threads multiply the places to wait by the number of connections, and multiplexing gathers the places to wait into one. Either way, the core is the same. You make one connection's wait not block another connection's progress.

What it looks like in the field

Incident reports come in this shape. Server CPU is idle, only the tail of the response time is long, and after a restart it gets better briefly and then worsens again. The health check passes. The side sending the health check is a diligent client that sends its request right away, so when its turn comes it gets an answer quickly.

The cause is usually a few slow clients. Request headers arrive late over a mobile line, a client opens a connection ahead of time and writes to it later, or someone maliciously sends one byte at a time. In a sequential server, one such client ties the throughput of the whole service to its own speed. If you look only at average latency, this situation is invisible. The time of the blocked clients hides behind the average.

What you will do in the next quiz

Draw it out yourself. If a slow client starts talking after 2 seconds and six fast clients arrive at 0.3 seconds, when do the six responses go out in a sequential server? When did their connect succeed? That these two are different moments is the core of this module. In the quiz, you check which calls stop the process, what the listen queue delays and what it cannot delay, and why raising the backlog is not a fix.