The heap had room, but the service stopped
OutOfMemoryError does not kill the process
Summary
When an OutOfMemoryError occurs, what you need is a heap dump of that moment and a setting that reliably kills the process. -XX:+HeapDumpOnOutOfMemoryError handles the former and -XX:+ExitOnOutOfMemoryError the latter. You find the culprit of a leak in the first line of the class histogram and in the dominator tree of the dump.
Why this was needed
There was a static HashMap named "cache." Things were put in and never removed. A day later the heap filled, Full GCs ran back to back, and then java.lang.OutOfMemoryError: Java heap space occurred. But the process did not die — the exception occurred in just one thread, and the other threads held on for hours in a state where they could not respond, repeating GC on almost no heap. The health check passed (that request uses almost no memory). The load balancer did not remove this instance. A restart brought it back, and with no dump the cause became "we will look next time it happens."
How it works
The java command documentation describes -XX:+HeapDumpOnOutOfMemoryError as "dumps the heap in HPROF format to the current directory when an OutOfMemoryError is thrown; off by default," and says -XX:HeapDumpPath=<경로> (the placeholder being the path) sets the file location and name, with the default name java_pid<pid>.hprof. In production you always turn these two on. A dump is created only at the moment the exception occurs, so if you do not turn it on, there is nothing after the incident. The file size is similar to the heap size, so reserve disk space in advance (the -Xmx64m lab produces about 50MB, measured).
-XX:+ExitOnOutOfMemoryError solves the problem of the process not dying. This option is not listed in the JDK 21 java command documentation, but HotSpot accepts it, and in measurement, at the first OutOfMemoryError it prints Terminating due to java.lang.OutOfMemoryError: Java heap space and ends with exit code 3. On Kubernetes the container dies and restarts, so the "hanging on half dead" state disappears. The meaning of this option is to turn a state that health checks cannot catch into a process exit.
On a live process you use two jcmd commands. GC.class_histogram produces heap usage statistics per class (impact: high — proportional to heap size). The output is rank, instance count, bytes, and class name, and a leak is almost always in the first line — if [B (byte arrays) or java.util.HashMap$Node number in the millions, the next question is who is holding them. GC.heap_dump <파일> creates an HPROF dump (the placeholder is the output file), and the documentation says that unless you give -all it first requests a Full GC — so the dump keeps only reachable objects, and that is the definition of "being held." An HPROF file starts with JAVA PROFILE 1.0.2 (measured). Analysis tools (Eclipse MAT and others) compute from this file the dominator tree — how much is freed if you let go of an object — and if a collection linked from a static field is at its top, that is the culprit.
Making the same allocations without holding the references is not a leak. The lab's Leak creates the same arrays with -Dleak.retain=false but does not put them in the map — the allocation volume is the same, yet it runs to the end even with -Xmx64m. The definition of a leak is not "creates a lot" but "does not let go."
Inside a container, the default heap comes from the cgroup limit. java -XX:+PrintFlagsFinal -version prints the actual MaxHeapSize, and if you change -XX:MaxRAMPercentage (default 25%), that value changes accordingly. On a Pod with a 2Gi memory limit, the default maximum heap is 512MB, and raising it to 50% makes it 1GiB. If you set the heap too close to the limit, the process dies not with OutOfMemoryError but as OOMKilled because of memory outside the heap (metaspace, thread stacks, direct buffers) — these are different incidents, and no dump is left either.
What it looks like in the field
The most common case is having no dump after the incident — the option was not turned on, or it was on but there was no disk, or the file vanished when the container restarted. The dump path must be a volume that survives. The second is leaving a "zombie that does not die" for hours without ExitOnOutOfMemoryError. The third is looking at the first line of the histogram and concluding "byte arrays are the problem." [B is always first. The question is who is holding them, and the answer is in the dominator tree of the dump. Finally, there is keeping startup options in a person's memory. Write the options in a startup script and keep that script in the repository.
What you will do in the next lab
You run Leak.java with -Xmx64m to create an OutOfMemoryError, get a dump with HeapDumpOnOutOfMemoryError, take a histogram and dump from a live process with jcmd, read the first line of the histogram, confirm with retain=false that the same allocation is not a leak, catch exit code 3 of ExitOnOutOfMemoryError, measure the container heap default with MaxRAMPercentage, and finally write a startup script containing all these options and actually run it.