Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
When the connection pool runs dry, fast requests queue too
In one line
When a connection pool dries up, the server is idle but all of the client's requests queue up, and the cause is usually not the pool size but the timeout and discard rules.
Why this was needed
HTTP clients, DB drivers, and gRPC channels all have a pool inside. It is meant to avoid paying the cost of making a connection — the TCP handshake, the TLS handshake, authentication — on every request. A pool usually runs quietly and well, and then one day collapses all at once. And the way it collapses looks far from the cause. Server CPU is idle and the server-side latency metrics are fine, but only the client's requests take several seconds, or "could not get a connection from the pool" errors pour out.
Little's Law explains this phenomenon in one line. The number of connections needed at the same time is the requests per second multiplied by the latency of one request. If you handle 200 per second at 50ms each, 10 connections are enough. If one service behind becomes slow and latency goes to 5 seconds, the same load needs 1,000 connections. If the pool limit is 10, the rest queue up. The pool running dry is a result, and the cause is latency.
How it works
There are five numbers that make up a pool, and each prevents a different incident.
| Setting | What it prevents | Without it |
|---|---|---|
| Maximum size | Opening connections to the server without limit | During an outage the client hammers the server |
| Acquire wait time | Waiting endlessly for a free slot | All calling threads stop in front of the pool |
| Request timeout | A slow call occupying a slot indefinitely | A few slow calls hold the whole pool |
| Idle limit | Reusing a connection the other side has already closed | After a quiet while, the first request fails with a reset |
| Maximum lifetime | Old connections piling onto one server | Even if you add servers, the load does not move over |
Acquire wait time and request timeout are often confused. The acquire wait time only rescues the side that is waiting; it cannot end the side that holds a slot. If four calls that take 5 seconds fill a pool of size 4, every fast call with an acquire wait time of 1.5 seconds fails. The slow calls must have a request timeout for the slots to come back.
A quieter trap is putting a timed-out connection back into the pool. A timeout means you stopped waiting, not that the server will not send a response. The late response arrives on the same socket, and the request that borrows that socket next reads someone else's response as its own. In a protocol like HTTP/1.1, where requests and responses are paired only by order, this becomes an incident where data gets mixed up. The rule is one. A connection whose state you do not know gets discarded.
The idle limit has to be shorter than the other side's idle limit. Servers and load balancers cut quiet connections first. The default keepAliveTimeout of the Node.js HTTP server is 5 seconds, so a client pool with a longer idle limit takes out a connection that rested for 6 seconds and gets a connection reset. The default idle limit of an AWS Application Load Balancer is 60 seconds.
Maximum lifetime has to do with load distribution. A pool tries to use a connection for a long time once it has made it, so even if you add servers, existing connections stay on the original servers. If you give connections a lifetime and close them when used up so they are made anew, the new connections go to the new servers, and the load spreads out evenly over time. To keep all connections from ending at the same moment, mix a little randomness into the lifetime too.
What it looks like in the field
The most common scene is that when one external API slows down, even a completely unrelated API slows down. The two were using the same HTTP client and the same pool, and the slow one took all of it. The remedy is to split the pool per destination (bulkhead), or at least to set the request timeout separately per destination.
The shape of the pool is a little different in HTTP/2 and gRPC clients. One connection carries several streams, so what you "borrow" is not a connection but a stream, and the upper bound is decided by the number of concurrent streams the other side announced (SETTINGS_MAX_CONCURRENT_STREAMS). When you hit that limit, new calls queue up inside the client, and this is invisible in the server metrics too. A service that tries to get by on one connection and gets blocked by the limit sometimes resolves it by opening a few more connections. In the end, you ask the same question again at every layer — what can happen at the same time, up to how many, and where does it wait when it overflows.
Observation also has to happen at the pool. The queue is invisible in server-side metrics, so the client has to export the pool's counts of connections in use, idle, and waiting, the acquire wait time, and the number of acquire failures. Without these numbers, all that gets said during an incident is "the server is fine." In particular, the acquire wait time tends to get reported mixed into request latency, so you have to measure it separately to tell whether the slowdown is the server or the queue in front of the pool.
What you will do in the next lab
You build a pool with a limit, a wait time, discard rules, metrics, and an idle limit using the standard library. You reproduce the scene where putting a timed-out socket back makes the next request read a stale response, and you measure with the same measuring tool the scene where four 5-second calls dry up the pool and the scene where a single 0.5-second timeout brings it back to life.