TT Lab
Get started
Learn Learning paths Courses

One Slow Connection Froze Every Other One

Where this design falls over

Continue in TT Lab

In one line

The number of connections one process can bear is decided by the file descriptor limit, the memory per connection, and a single core, and you can measure and confirm all three.

Why this was needed

Once you switch to multiplexing, it looks like you can hold up even when the connection count suddenly grows. So you stop asking where the limit is, and the limit is usually discovered for the first time on the busiest day. This module splits that limit into three and counts it in advance. All three are values you can confirm with a single command, not guesses.

How it works

The first is the file descriptor limit. According to getrlimit(2), RLIMIT_NOFILE means a value one greater than the highest descriptor number that process can open, and an attempt to exceed it fails with EMFILE. One connection is one descriptor, so this value is the ceiling on concurrent connections. You can see your own process's value with ulimit -n, and count /proc/<pid>/fd to see how many it actually holds.

What gets counted here is not only connections. There are the three standard I/O streams, one listening socket, and, if the implementation of the multiplexing object uses a descriptor, one more. So if you attach 200 connections and count, the value is not 200 but slightly more. The reason this small difference matters is that only if you can explain it can you predict what runs out first near the limit.

The second is the limit of the watching mechanism itself. As seen earlier, select(2) can handle only numbers below FD_SETSIZE, which is 1024, in the glibc implementation. However high you raise the descriptor limit, as long as you use select you cannot watch descriptors past number 1023, and the document says to use poll or epoll in that case. If you use Python's DefaultSelector, the platform makes this choice for you, but it is worth checking which implementation was picked.

The third is the memory per connection. It includes not only the state the program created but also the kernel's socket buffers. tcp(7) explains that the socket buffer size is set through /proc/sys/net/ipv4/tcp_rmem and tcp_wmem, or per socket through SO_RCVBUF and SO_SNDBUF, and says that TCP actually allocates twice the requested size and uses the extra space for bookkeeping structures. That is why the value you read back with getsockopt is not the same as the value you set. If you leave out this multiplier when computing the room one connection takes, the budget is off by half.

Limit Where it is decided How to check
Number of concurrent connections RLIMIT_NOFILE ulimit -n, counting /proc/PID/fd
Numbers you can watch FD_SETSIZE of select 1024 per the document; not applicable to epoll
Memory per connection Socket buffers and program state tcp_rmem and actual usage

There is a fourth on top of this. An event loop has a single flow of execution, so it uses only one core. When connections grow and CPU fills one core, latency alone grows even though descriptors and memory remain. The solution then is not to optimize the loop but to add processes and use more cores. That is why real services use multiplexing and multiple processes together. This course covers the first half of that.

What it looks like in the field

Cases where the same code works on a laptop but not in a container often come from this limit. The descriptor limit is a per-process value, so it changes when the runtime environment changes, and when you hit the limit, the EMFILE that comes out usually blows up not at the place that accepts connections but somewhere unrelated. Code that opens files or rotates logs fails first, so the fact that the cause is the network is discovered late.

Memory incidents are quieter. A few kilobytes per connection looks small, but with tens of thousands it becomes gigabytes, and since this room is on the kernel side rather than in the application heap, it hardly shows up in the language runtime's memory metrics. So you get a situation where the process's resident memory is ordinary but the node is under pressure.

So when you make a capacity plan, decide the target connection count first, and multiply that count by the descriptor limit and by the room per connection separately. Whichever is hit first is the service's real limit. If you write the numbers down, it becomes clear what to increase when you hit the limit later, and above all you can know before you hit it.

What you will do in the next lab

Now you will build it yourself. First you measure the queueing with a sequential server, then receive the same load with a multiplexing server and compare the two numbers. At the end, you attach 200 connections, count the descriptors the server process holds, and you should be able to explain why that value is not 200. The lab's grader re-measures the numbers you write down on the spot and compares them, so writing a plausible value will not pass.