TT Lab
Get started
Learn Learning paths Courses

Real-Time Communication — WebSocket, gRPC Streaming and WebRTC

Build a connection pool and drain it

Continue in TT Lab

Goal

Build yourself a client connection pool with a limit, a wait time, discard rules, and an idle limit, and measure with the same measuring tool the scene where slow calls with no request timeout dry up the pool and the scene where it comes back to life.

Why it matters

HTTP clients, DB drivers, and gRPC channels all have a pool inside. When the pool dries up, the server is idle but all of the client's requests queue up, and to someone looking only at server metrics, nothing is visible. The cause is usually not the pool settings but the rules outside the pool — there is no request timeout, a timed-out connection is put back, or a connection the other side has already closed is reused. This lab reproduces and blocks those three one by one. It uses only the standard library.

Steps

  1. A pool with a limit — In /root/rt/pool/pool.py, create the PoolTimeout exception and the Pool class. Pool(factory, max_size, acquire_timeout, idle_timeout=None) takes a factory that makes one connection when called with no arguments. acquire() returns an idle connection if there is one, and if there is none and the number of connections made so far is less than max_size, it makes a new one with the factory and returns it. release(conn) puts the connection back in the idle list. It is a ValueError if max_size is not an int of 1 or more or if acquire_timeout is not positive, and a bool is not accepted as a number.
  2. If there is no free slot, wait only as long as set — Fix acquire() so that if all connections are in use and no more can be made, it waits up to acquire_timeout seconds. If someone calls release in the meantime, the waiting side receives that connection right away, and when the time is up, it raises PoolTimeout. Do not spin using the CPU while waiting.
  3. Do not put back a connection whose state you do not know — Add the broken argument of release(conn, broken=False) to Pool and add the connection() context manager. If broken=True, call the connection's close() and remove it from the pool, so the next acquire can make a new connection. If the with pool.connection() as conn: block ends normally, put the connection back, and if it ends with an exception, treat it as broken and then re-raise that exception as it is.
  4. Put out the pool's state as numbers — Add stats() to Pool. Return a dict with five keys: size (the number of connections currently alive), idle (the number of idle connections), in_use (the number of borrowed connections), waiting (the number of calls currently waiting in acquire), and timeouts (the number of times PoolTimeout has been raised so far).
  5. A timed-out connection still has the previous response on it — In /root/rt/pool/pool.py, add get(pool, path, timeout). Borrow a connection (a socket) from the pool, apply settimeout(timeout), send an HTTP/1.1 GET with h1_get(sock, path) from /opt/fixtures/rt/rtnet.py, and return (status, body). If any exception occurs, including a timeout, discard that connection as broken and re-raise the exception. The grader times out a slow request with a pool of max_size 1 and then immediately sends the next request.
  6. Four slow calls hold the whole pool — Run /opt/rt-lab/bin/python /opt/fixtures/rt/pool/starve.py. With your pool of size 4, it first sends four requests that take 5 seconds and then sends twenty fast requests. It prints how many fast requests succeeded with no request timeout and with 0.5 seconds. Write the two output lines starved_ok and bounded_ok into /root/rt/pool/report.txt. The grader re-measures and compares.
  7. Do not reuse a connection the server closed first — Implement Pool's idle_timeout. Record the time a connection came back through release, and when acquire takes out an idle connection, one whose idle time exceeds idle_timeout seconds is closed with close() and discarded, and then the next one is checked. If it is None, nothing is discarded. The grader puts the server behind a relay that cuts connections quiet for 1 second, and with a pool of idle_timeout 0.5, sends a request, rests 1.5 seconds, and sends another request.

Notes

A pool with a limit

In /root/rt/pool/pool.py, create the PoolTimeout exception and the Pool class. Pool(factory, max_size, acquire_timeout, idle_timeout=None) takes a factory that makes one connection when called with no arguments. acquire() returns an idle connection if there is one, and if there is none and the number of connections made so far is less than max_size, it makes a new one with the factory and returns it. release(conn) puts the connection back in the idle list. It is a ValueError if max_size is not an int of 1 or more or if acquire_timeout is not positive, and a bool is not accepted as a number.

Using idle connections first is the reason a pool exists. It is meant to avoid paying the cost of making a connection (TCP handshake, TLS) on every request. The grader counts how many times the factory was called.

If there is no free slot, wait only as long as set

Fix acquire() so that if all connections are in use and no more can be made, it waits up to acquire_timeout seconds. If someone calls release in the meantime, the waiting side receives that connection right away, and when the time is up, it raises PoolTimeout. Do not spin using the CPU while waiting.

threading.Condition's wait(timeout) does not tell you why it woke up. Each time it wakes, you have to check the condition again and recompute the remaining time. If you fix the deadline once with time.monotonic(), the calculation gets simple.

Do not put back a connection whose state you do not know

Add the broken argument of release(conn, broken=False) to Pool and add the connection() context manager. If broken=True, call the connection's close() and remove it from the pool, so the next acquire can make a new connection. If the with pool.connection() as conn: block ends normally, put the connection back, and if it ends with an exception, treat it as broken and then re-raise that exception as it is.

At the moment an exception occurs, you do not know what remains on that connection. The rest of a response may still be on the socket. Discarding what you do not know is the pool's basic rule. You can write it briefly with contextlib.contextmanager.

Put out the pool's state as numbers

Add stats() to Pool. Return a dict with five keys: size (the number of connections currently alive), idle (the number of idle connections), in_use (the number of borrowed connections), waiting (the number of calls currently waiting in acquire), and timeouts (the number of times PoolTimeout has been raised so far).

Pool exhaustion is not visible in server metrics. That is because the server is idle and only the client is queued up. waiting and timeouts are the only numbers that show that queue. Raise waiting when waiting starts and lower it when the wait ends for any reason.

A timed-out connection still has the previous response on it

In /root/rt/pool/pool.py, add get(pool, path, timeout). Borrow a connection (a socket) from the pool, apply settimeout(timeout), send an HTTP/1.1 GET with h1_get(sock, path) from /opt/fixtures/rt/rtnet.py, and return (status, body). If any exception occurs, including a timeout, discard that connection as broken and re-raise the exception. The grader times out a slow request with a pool of max_size 1 and then immediately sends the next request.

A timeout means you stopped waiting, not that the server will not send a response. That response arrives on the same socket, even if late. If you put that socket back in the pool, the next request reads someone else's response. If you put /opt/fixtures/rt in sys.path, you can import and use rtnet.

Four slow calls hold the whole pool

Run /opt/rt-lab/bin/python /opt/fixtures/rt/pool/starve.py. With your pool of size 4, it first sends four requests that take 5 seconds and then sends twenty fast requests. It prints how many fast requests succeeded with no request timeout and with 0.5 seconds. Write the two output lines starved_ok and bounded_ok into /root/rt/pool/report.txt. The grader re-measures and compares.

A call with no timeout occupies a slot in the pool indefinitely. The acquire_timeout only rescues the side that waits; it cannot end the side that holds a slot. You should be able to explain why the two values come out that way.

Do not reuse a connection the server closed first

Implement Pool's idle_timeout. Record the time a connection came back through release, and when acquire takes out an idle connection, one whose idle time exceeds idle_timeout seconds is closed with close() and discarded, and then the next one is checked. If it is None, nothing is discarded. The grader puts the server behind a relay that cuts connections quiet for 1 second, and with a pool of idle_timeout 0.5, sends a request, rests 1.5 seconds, and sends another request.

Load balancers and servers cut quiet connections first. A client pool learns that only the next time it sends on that connection. If you set the pool's idle limit shorter than the other side's idle limit, you avoid that race.