What maxThreads 200 Actually Means
In one line
acceptCount, maxConnections and maxThreads each do a different job at a different point, and when the cause of slow responses is not a shortage of threads, raising maxThreads makes the situation worse.
Why this distinction is needed
If you lump the three values together, the remedy is always the same — raise the number. But each blocks at a different point, so the symptoms differ too. When acceptCount is exceeded, the kernel refuses the connection itself and you get Connection refused; when maxConnections is exceeded, the connection is made but no response comes; when maxThreads is exceeded, requests wait their turn in the queue. You should be able to pick which number to touch from the symptom alone.
The three numbers each do different work
When a performance issue arises, someone always says, "Let's raise maxThreads." But in most cases that action changes nothing, or makes things worse. To know why, you need to know at which point each of the three values operates.
[클라이언트 소켓]
|
v
(1) acceptCount ← OS 소켓 백로그 큐. 여기까지 넘치면 Connection refused
|
v
(2) maxConnections ← 톰캣이 동시에 '들고 있는' 커넥션 수
|
v
(3) maxThreads ← 동시에 '처리 중인' 요청 수 (워커 스레드)
|
v
[애플리케이션]
- maxThreads (default 200) — the number of worker threads processing requests at the same time. One request occupies one thread. It occupies it even while waiting for the DB.
- maxConnections (NIO default 10000) — the number of sockets Tomcat keeps. A connection that is merely open through Keep-Alive and sends no request does not eat a thread, so this value can be much larger than maxThreads. That is why NIO is used.
- acceptCount (default 100) — the number of connections waiting in the OS socket backlog
when even maxConnections is full. When this overflows too, the client gets not a 503 but
Connection refused. This difference is decisive in incident analysis.
A Connection refused in the log may mean not "the server is dead" but
"the server is alive but had no capacity to accept". If the nginx log in front shows
connect() failed (111: Connection refused), you must put backlog saturation on the list of candidates
as well as the death of the backend process.
Why raising maxThreads doesn't work
If most of the request processing time is DB waiting, adding threads is adding people queuing in front of the DB. Throughput stays the same and only the waiting time grows. In bad cases the connection pool is exhausted and everything stops.
If you take a thread dump and see this pattern, it is not a thread pool problem.
"http-nio-8080-exec-73" ... WAITING (parking)
at jdk.internal.misc.Unsafe.park
at com.zaxxer.hikari.pool.HikariPool.getConnection ← DB 커넥션 풀 대기
The message of this dump is clear: the problem is not the WAS thread pool but the DB connection pool. If you raise maxThreads to 400, the waiting threads merely increase from 200 to 400.
Sizing with evidence — Little's law
Don't decide by feel; calculate. Little's law is written like this.
동시 처리 수 = 목표 처리량(TPS) × 평균 응답시간(초)
If the target is 300 per second and the average response is 200ms,
300 × 0.2 = 60
that is, at least 60 threads must be working at the same time to get 300 TPS.
You multiply this by a headroom rate (relative to peak, usually 20–50%). With 30% headroom, 60 × 1.3 = 78.
For CPU-bound work it is different. Threads beyond the core count only add context-switching cost. The rule of thumb is roughly this.
- I/O-bound (most business web apps):
코어 수 × (1 + 대기시간 / 처리시간) - CPU-bound (encryption, image processing):
코어 수 + 1
And what you must check — after you decide maxThreads, look at its relationship with the DB connection pool.
If 200 WAS threads all use the DB and the pool is 20, 180 wait.
Conversely, if you enlarge the pool to 200, this time the DB's max_connections blows up.
With several instances, it becomes a multiplication.
서비스 A 12대 × 풀 20 = 240
서비스 B 8대 × 풀 15 = 120
배치 워커 4대 × 풀 5 = 20
합계 = 380
DB max_connections = 200 ← 롤링 배포로 인스턴스가 잠깐 두 배 되는 순간 터진다
A system that hasn't done this calculation collapses during a rolling deployment with
FATAL: sorry, too many clients already. Worse, that error
also blocks health checks and monitoring agents. At the moment of the outage, the means of observation disappear together.
Keep-Alive cuts both ways
<Connector port="8080" protocol="HTTP/1.1"
connectionTimeout="20000"
keepAliveTimeout="5000"
maxKeepAliveRequests="100"
maxThreads="150" minSpareThreads="20"
maxConnections="2000" acceptCount="50" />
connectionTimeout— after accepting a connection, the time it waits until the request line arrives. It is not a limit on request processing time. Confusing this produces "I increased the timeout but it's still slow".keepAliveTimeout— the time it keeps a connection while waiting for the next request. If long, connection reuse reduces latency, but idle connections eat into maxConnections.maxKeepAliveRequests— the maximum number of requests handled on one connection.-1is unlimited.
If nginx is in front, the connections between nginx and Tomcat are few and heavily reused. In that case it is better to give keepAliveTimeout generously. Conversely, if Tomcat is exposed directly to the internet, idle connections eat resources, so keep it short. It is not a "correct value" but "a value that depends on the front-end configuration".
A shared Executor
If you have several connectors, each has its own thread pool. Then it is hard to control the total number of threads. In that case you use a shared Executor.
<Executor name="tomcatThreadPool" namePrefix="labhub-exec-"
maxThreads="150" minSpareThreads="20"/>
<Connector executor="tomcatThreadPool" port="8080" protocol="HTTP/1.1" ... />
Giving namePrefix a meaningful value is a practical tip. When you take a thread dump,
if it shows labhub-exec-37, you know at once which connector's thread it is.
The default http-nio-8080-exec- isn't bad either, but when there are several connectors you need to tell them apart.
Turn on heap and GC logs before tuning
CATALINA_OPTS="-Xms1g -Xmx1g \
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/logs/dump \
-Xlog:gc*:file=/logs/gc.log:time,uptime:filecount=5,filesize=10m"
- Setting
-Xmsand-Xmxequal is the convention for server applications. It eliminates the GC caused by the heap growing and shrinking. -XX:+HeapDumpOnOutOfMemoryErroris not optional but mandatory. OOMs are hard to reproduce. If you can't get a dump at the moment it blows, root cause analysis becomes guesswork. However, the dump file is as large as the heap, so check that you have free disk space.- GC logs are too late if you turn them on after a performance problem arises. Turn them on from the start and set file rotation.
If you put these three lines in before go-live, the time for incident analysis in the stabilization period is halved.
What you see in the field
The remedy that comes up most often in performance meetings is "let's raise maxThreads". And usually nothing gets better. Threads remain occupied while requests wait for DB responses, so if the bottleneck is the DB, adding threads only increases the number of waiting threads. If anything, context switching and heap usage increase and GC pauses get longer.
So there is an order. First take a thread dump and see what the workers are doing. If most are RUNNABLE, the CPU is short; if most are WAITING on sockets or locks, the bottleneck is not the threads but what is behind them. In the latter case, raising maxThreads is enlarging the waiting room, not adding counters.