The heap had room, but the service stopped
Find why it stopped with a thread dump
Goal
Reproduce "the heap has room but the service stopped" yourself and pinpoint the cause with thread dumps — the thread holding the lock, the queued threads, deadlock, and recovery using a tryLock timeout.
Why it matters
When the process and port are alive but there is no response, the heap graph tells you nothing. Only a thread dump shows "who is waiting for what." A restart destroys the evidence, so the hand that takes a dump at the moment of the stall comes first. The material Stuck.java is an HTTP service running on 127.0.0.1:8085 with 4 worker threads (worker-N) — /health does not touch the DB, / takes the DB lock and works for 5ms, and /poison takes the lock and sleeps forever. If you pass -Ddb.lock.timeout.ms=<ms>, it replaces synchronized with ReentrantLock.tryLock and returns 503 if it cannot get the lock within that time.
Steps
- Compile Stuck.java in
/root/jvm/threads, start it withnohup java -cp /root/jvm/threads Stuck > /root/jvm/threads/server.log 2>&1 &, and write the pid to/root/jvm/threads/server.pid.curl -s http://127.0.0.1:8085/healthmust return ok. - Save the list of JVMs with
jcmd -l > /root/jvm/threads/jcmd.txt. Stuck must appear with the pid from server.pid. - Save a dump of the normal state with
jcmd <pid> Thread.print > /root/jvm/threads/dump-idle.txt. It must containFull thread dumpand"worker-1". - Create the incident: send
curl -s -m 20 http://127.0.0.1:8085/poison &to make it hold the lock forever, then sendcurl -s -m 20 http://127.0.0.1:8085/ &three more times (all in the background). After 2 seconds, runjcmd <pid> Thread.print > /root/jvm/threads/dump-stuck.txt. The dump must have at least 3waiting to locklines andBLOCKED, and/healthalso stops responding. - From dump-stuck.txt, read the name of the thread holding
- locked <주소> (a Stuck$Db)and the number of threads thatwaiting to lockthe same address, and write them to/root/jvm/threads/diagnosis.txtas two lines,holder=<스레드이름>(the thread name) andblocked=<n>. The grader compares against a dump of the live process. - Compile
Deadlock.java, start it withjava -cp /root/jvm/threads Deadlock > /root/jvm/threads/deadlock.log 2>&1 &(leave it running), and after 1 second runjcmd <pid> Thread.print > /root/jvm/threads/dump-deadlock.txt. It must containFound 1 deadlock.. - Recovery: kill the process in server.pid, restart it with
-Ddb.lock.timeout.ms=500, and update server.pid. Send/poisonin the background again, then save the result ofcurl -s -m 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:8085/to/root/jvm/threads/after.txt— now a 503 must come back within 5 seconds. - With the service still running, save
/root/jvm/threads/dumps/dump-1.txt,dump-2.txt, anddump-3.txtat 2-second intervals. The time lines of the three files (the second line after the pid line) must all differ.
Notes
- Finding the lock address:
grep -n 'locked <.*Stuck\$Db' dump-stuck.txtandgrep -c 'waiting to lock <주소>' dump-stuck.txt. The thread name is the nearest line above that line that starts with"...". - If you start it twice on the same port, the second one dies from a bind failure. Kill the first one in step 7 first.
- Common mistakes: taking only one dump, and counting only the queued threads without finding the thread holding the lock.
Start the service and save the pid
Compile Stuck.java in /root/jvm/threads, start it with nohup, and write the pid to /root/jvm/threads/server.pid. /health returns ok.
After nohup java -cp /root/jvm/threads Stuck > server.log 2>&1 &, run echo $! > server.pid. If you chain with && as in cd … && nohup … &, $! becomes the pid of a subshell, so do the cd separately. Startup takes about a second, so wait briefly before curl.
The JVM list
Save jcmd -l > /root/jvm/threads/jcmd.txt. Stuck must appear with the pid from server.pid.
jcmd -l shows only JVMs started by the same user. Each line is a pid and a main class.
A dump of the normal state
Save jcmd Thread.print > /root/jvm/threads/dump-idle.txt. It must contain Full thread dump and "worker-1".
You can pull out the pid with $(cat /root/jvm/threads/server.pid). Save the normal dump first so you can compare it with the incident dump.
Create the incident and take a dump
Send /poison in the background, send / three more times in the background, and after 2 seconds run jcmd Thread.print > /root/jvm/threads/dump-stuck.txt. It must have at least 3 waiting to lock lines and BLOCKED.
Run curl -s -m 20 http://127.0.0.1:8085/poison &, then curl -s -m 20 http://127.0.0.1:8085/ & three times. When all 4 worker threads stand in front of the lock, /health stops responding too — that is the incident.
The thread holding the lock and the queued threads
From dump-stuck.txt, write the name of the thread holding locked (a Stuck$Db) and the number of threads that waiting to lock the same address to /root/jvm/threads/diagnosis.txt as holder= and blocked=.
Take the address (<0x...>) from the locked line; the nearest line above it that starts with a double quote is the thread name. blocked is the number of waiting to lock lines with the same address.
The JVM names a deadlock for you
Compile Deadlock.java, start it in the background (leave it running), and run jcmd Thread.print > /root/jvm/threads/dump-deadlock.txt. It must contain Found 1 deadlock.
After java -cp /root/jvm/threads Deadlock > deadlock.log 2>&1 &, the pid is $!. The two threads wait on each other's locks, so the program does not end — leave it until grading is finished.
Turn the infinite wait into a 503
Kill the process in server.pid, restart it with -Ddb.lock.timeout.ms=500, and update server.pid. After sending /poison in the background, save the result of curl -s -m 5 -o /dev/null -w '%{http_code}' http://127.0.0.1:8085/ to /root/jvm/threads/after.txt (503).
System properties come before the class name: java -Ddb.lock.timeout.ms=500 -cp ... Stuck. If tryLock cannot get the lock within 500ms it returns 503, so the service only slows down; it does not stop.
Three dumps
With the service running, save /root/jvm/threads/dumps/dump-1.txt, dump-2.txt, and dump-3.txt at 2-second intervals. The time lines of the three files (the second line after the pid line) must all differ.
for i in 1 2 3; do jcmd $PID Thread.print > dumps/dump-$i.txt; sleep 2; done. If the same thread is in the same place in all three dumps, it really is stuck.