TT Lab
Get started
Learn Learning paths Courses

The heap had room, but the service stopped

A service stops at its smallest pool

Continue in TT Lab

Summary

A service stops at its smallest pool. Even with 200 request-handling threads, if the DB connection pool has 10 connections and waiting has no upper bound, a single slow query lines 190 threads up in front of the connections. Pool size and the upper bound on wait time must be decided as a pair.

Why this was needed

The previous module's incident was a lock. What you see more often in the field is pool exhaustion. The DB slowed down for a moment. All 10 connections in the pool were held by slow queries. The following requests wait for a connection indefinitely. Once all 200 request-handling threads are in that wait, there is no thread left even to handle a health-check request. Even if the DB returns to normal a minute later, the service stays dead until the backed-up requests have all drained. In the meantime the load balancer removes this instance, load shifts to the remaining instances, and the same thing spreads. The problem is not the pool size but that waiting has no upper bound.

How it works

Tomcat's HTTP connector documentation explains the path a request takes with three numbers. maxThreads — the maximum number of request-handling threads, that is, the number of requests that can be handled concurrently; the default is 200. maxConnections — the number of connections the server can accept and hold as in progress; when this number is reached, connections are accepted but not processed, and wait. acceptCount — the number of connection requests the operating system queues when maxConnections is full; the default is 100; once even this queue is full, the operating system refuses the connection or times it out. In other words, requests get pushed back in the order thread → connection → OS queue. minSpareThreads is the number of threads always kept alive, with a default of 10. connectionTimeout is the number of milliseconds to wait for the request line after accepting a connection; the documentation default is 60000, but the distribution's server.xml sets 20000. keepAliveTimeout is the time to wait for the next request, and if not specified it follows connectionTimeout. If you use the executor attribute to attach a shared <Executor> to the connector, these thread attributes are ignored and the Executor's values are used (the tc-threadpool lab is that path).

Thread names are the key to diagnosis. The connector's internal pool threads are named http-nio-8080-exec-N, and counting this prefix in a thread dump immediately shows how many are alive right now and how many of them are waiting where.

For the DB connection pool, the DBCP 2 example in the JNDI datasource documentation is the reference. In context.xml, in <Resource type="javax.sql.DataSource" …>, you set maxTotal (the maximum number of connections in the pool; -1 means unlimited), maxIdle (the maximum number to keep idle), and maxWaitMillis — the maximum milliseconds to wait until a connection is free; past it an exception is thrown, and -1 means wait indefinitely. The cause of the incident is exactly this -1. With maxWaitMillis="2000", an exception occurs after 2 seconds, and the application can turn it into a 503 and return the thread.

Seen in Java code, the same principle is that Semaphore(N) is the connection pool. acquire() waits indefinitely, while tryAcquire(timeout, unit) sets an upper bound. In a dump, an acquire() wait appears as parking to wait for <…> (a java.util.concurrent.Semaphore$FairSync). The lab's PoolDemo builds exactly this shape — 4 worker threads, 2 connections, 3-second queries. If you send 6 requests at once, 2 work and the rest stand in front of the connections. If you give -Ddb.pool.timeout.ms=500, a 503 comes back after 500ms.

Putting the management port on a different thread is also a design decision. The /stats of PoolDemo is answered by a separate thread on 8087, so you can read the state even when all worker threads are blocked. This is the same reason Tomcat has JMX or a separate connector — leave a way to ask a stalled service.

The convention for running Tomcat in multiple copies is CATALINA_HOME and CATALINA_BASE from the introduction document. HOME is the installation (bin, lib), and BASE is the per-instance configuration, logs, and web apps (conf, logs, temp, webapps, work). If you change configuration by creating only a new BASE without touching the installation, you can isolate an experimental instance from the production configuration. You start it that way in the lab.

What it looks like in the field

The most common response is to only increase the size. If you raise maxThreads from 200 to 800, 800 threads stand in front of the connections — it stalls longer and bigger. The answer is to size the pool only as much as the downstream (the DB) can handle, and to put an upper bound on waiting so requests fail fast. The second is a health check that goes through the DB. When the DB slows down, the health check fails too, and even healthy instances are removed. A health check must be a path that does not touch the pool. The third is putting a timeout on only one layer. Connection wait, query execution, HTTP client, and load balancer — every layer needs an upper bound, and the outer ones must be longer than the inner ones.

What you will do in the next lab

You start PoolDemo.java, drain the connection pool with concurrent requests, and confirm the Semaphore wait in a dump, then restart with -Ddb.pool.timeout.ms=500 and see a 503 come back. Next, you create a new CATALINA_BASE at /root/jvm/pool/tc, configure the 8080 connector in server.xml with maxThreads, acceptCount, and connectionTimeout, add to context.xml a DataSource containing maxTotal, maxIdle, and maxWaitMillis, actually

start that instance, and count the http-nio-8080-exec- threads in a dump.