Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
Build a connection pool and drain it
Goal
Build yourself a client connection pool with a limit, a wait time, discard rules, and an idle limit, and measure with the same measuring tool the scene where slow calls with no request timeout dry up the pool and the scene where it comes back to life.
Why it matters
HTTP clients, DB drivers, and gRPC channels all have a pool inside. When the pool dries up, the server is idle but all of the client's requests queue up, and to someone looking only at server metrics, nothing is visible. The cause is usually not the pool settings but the rules outside the pool — there is no request timeout, a timed-out connection is put back, or a connection the other side has already closed is reused. This lab reproduces and blocks those three one by one. It uses only the standard library.
Steps
- A pool with a limit — In /root/rt/pool/pool.py, create the PoolTimeout exception and the Pool class. Pool(factory, max_size, acquire_timeout, idle_timeout=None) takes a factory that makes one connection when called with no arguments. acquire() returns an idle connection if there is one, and if there is none and the number of connections made so far is less than max_size, it makes a new one with the factory and returns it. release(conn) puts the connection back in the idle list. It is a ValueError if max_size is not an int of 1 or more or if acquire_timeout is not positive, and a bool is not accepted as a number.
- If there is no free slot, wait only as long as set — Fix acquire() so that if all connections are in use and no more can be made, it waits up to acquire_timeout seconds. If someone calls release in the meantime, the waiting side receives that connection right away, and when the time is up, it raises PoolTimeout. Do not spin using the CPU while waiting.
- Do not put back a connection whose state you do not know — Add the broken argument of release(conn, broken=False) to Pool and add the connection() context manager. If broken=True, call the connection's close() and remove it from the pool, so the next acquire can make a new connection. If the with pool.connection() as conn: block ends normally, put the connection back, and if it ends with an exception, treat it as broken and then re-raise that exception as it is.
- Put out the pool's state as numbers — Add stats() to Pool. Return a dict with five keys: size (the number of connections currently alive), idle (the number of idle connections), in_use (the number of borrowed connections), waiting (the number of calls currently waiting in acquire), and timeouts (the number of times PoolTimeout has been raised so far).
- A timed-out connection still has the previous response on it — In /root/rt/pool/pool.py, add get(pool, path, timeout). Borrow a connection (a socket) from the pool, apply settimeout(timeout), send an HTTP/1.1 GET with h1_get(sock, path) from /opt/fixtures/rt/rtnet.py, and return (status, body). If any exception occurs, including a timeout, discard that connection as broken and re-raise the exception. The grader times out a slow request with a pool of max_size 1 and then immediately sends the next request.
- Four slow calls hold the whole pool — Run /opt/rt-lab/bin/python /opt/fixtures/rt/pool/starve.py. With your pool of size 4, it first sends four requests that take 5 seconds and then sends twenty fast requests. It prints how many fast requests succeeded with no request timeout and with 0.5 seconds. Write the two output lines starved_ok and bounded_ok into /root/rt/pool/report.txt. The grader re-measures and compares.
- Do not reuse a connection the server closed first — Implement Pool's idle_timeout. Record the time a connection came back through release, and when acquire takes out an idle connection, one whose idle time exceeds idle_timeout seconds is closed with close() and discarded, and then the next one is checked. If it is None, nothing is discarded. The grader puts the server behind a relay that cuts connections quiet for 1 second, and with a pool of idle_timeout 0.5, sends a request, rests 1.5 seconds, and sends another request.
Notes
- The working folder is /root/rt/pool. Create it first with mkdir -p /root/rt/pool.
- The measuring tool is /opt/rt-lab/bin/python /opt/fixtures/rt/pool/starve.py, and it calls your Pool and get. The measuring tool starts the server itself.
- The grader calls acquire concurrently from several threads. Make all state changes inside a lock.
- Two common mistakes. Putting a timed-out socket back in the pool so the next request reads a stale response, and not rechecking the condition after waking from wait, so you hijack someone else's connection.
- Always run Python with /opt/rt-lab/bin/python. This lab's libraries are only in that virtual environment, and if you run it with plain python3, you get a ModuleNotFoundError. It is convenient to shorten it with something like alias rpy=/opt/rt-lab/bin/python.
- The lab Pod blocks outbound connections. All communication happens on 127.0.0.1 inside the same Pod, and no installation or download is needed.
- The grader loads your code in a separate process and makes real connections. The example file is only a function skeleton, so it does not pass if left as is. Do not delete the functions you finished in earlier steps.
- When the lab session ends, the files in /root do not remain. Keep the code you need separately before you finish.
A pool with a limit
In /root/rt/pool/pool.py, create the PoolTimeout exception and the Pool class. Pool(factory, max_size, acquire_timeout, idle_timeout=None) takes a factory that makes one connection when called with no arguments. acquire() returns an idle connection if there is one, and if there is none and the number of connections made so far is less than max_size, it makes a new one with the factory and returns it. release(conn) puts the connection back in the idle list. It is a ValueError if max_size is not an int of 1 or more or if acquire_timeout is not positive, and a bool is not accepted as a number.
Using idle connections first is the reason a pool exists. It is meant to avoid paying the cost of making a connection (TCP handshake, TLS) on every request. The grader counts how many times the factory was called.
If there is no free slot, wait only as long as set
Fix acquire() so that if all connections are in use and no more can be made, it waits up to acquire_timeout seconds. If someone calls release in the meantime, the waiting side receives that connection right away, and when the time is up, it raises PoolTimeout. Do not spin using the CPU while waiting.
threading.Condition's wait(timeout) does not tell you why it woke up. Each time it wakes, you have to check the condition again and recompute the remaining time. If you fix the deadline once with time.monotonic(), the calculation gets simple.
Do not put back a connection whose state you do not know
Add the broken argument of release(conn, broken=False) to Pool and add the connection() context manager. If broken=True, call the connection's close() and remove it from the pool, so the next acquire can make a new connection. If the with pool.connection() as conn: block ends normally, put the connection back, and if it ends with an exception, treat it as broken and then re-raise that exception as it is.
At the moment an exception occurs, you do not know what remains on that connection. The rest of a response may still be on the socket. Discarding what you do not know is the pool's basic rule. You can write it briefly with contextlib.contextmanager.
Put out the pool's state as numbers
Add stats() to Pool. Return a dict with five keys: size (the number of connections currently alive), idle (the number of idle connections), in_use (the number of borrowed connections), waiting (the number of calls currently waiting in acquire), and timeouts (the number of times PoolTimeout has been raised so far).
Pool exhaustion is not visible in server metrics. That is because the server is idle and only the client is queued up. waiting and timeouts are the only numbers that show that queue. Raise waiting when waiting starts and lower it when the wait ends for any reason.
A timed-out connection still has the previous response on it
In /root/rt/pool/pool.py, add get(pool, path, timeout). Borrow a connection (a socket) from the pool, apply settimeout(timeout), send an HTTP/1.1 GET with h1_get(sock, path) from /opt/fixtures/rt/rtnet.py, and return (status, body). If any exception occurs, including a timeout, discard that connection as broken and re-raise the exception. The grader times out a slow request with a pool of max_size 1 and then immediately sends the next request.
A timeout means you stopped waiting, not that the server will not send a response. That response arrives on the same socket, even if late. If you put that socket back in the pool, the next request reads someone else's response. If you put /opt/fixtures/rt in sys.path, you can import and use rtnet.
Four slow calls hold the whole pool
Run /opt/rt-lab/bin/python /opt/fixtures/rt/pool/starve.py. With your pool of size 4, it first sends four requests that take 5 seconds and then sends twenty fast requests. It prints how many fast requests succeeded with no request timeout and with 0.5 seconds. Write the two output lines starved_ok and bounded_ok into /root/rt/pool/report.txt. The grader re-measures and compares.
A call with no timeout occupies a slot in the pool indefinitely. The acquire_timeout only rescues the side that waits; it cannot end the side that holds a slot. You should be able to explain why the two values come out that way.
Do not reuse a connection the server closed first
Implement Pool's idle_timeout. Record the time a connection came back through release, and when acquire takes out an idle connection, one whose idle time exceeds idle_timeout seconds is closed with close() and discarded, and then the next one is checked. If it is None, nothing is discarded. The grader puts the server behind a relay that cuts connections quiet for 1 second, and with a pool of idle_timeout 0.5, sends a request, rests 1.5 seconds, and sends another request.
Load balancers and servers cut quiet connections first. A client pool learns that only the next time it sends on that connection. If you set the pool's idle limit shorter than the other side's idle limit, you avoid that race.