TT Lab
Get started
Learn Learning paths Courses

Linux Fundamentals

Signals: What You Can Catch and What You Cannot

Continue in TT Lab

In one sentence

A signal is a short notification sent to a process, and of these, only SIGKILL (9) and SIGSTOP (19) cannot be caught, blocked, or ignored by the program. For all the others, the application can change how they are handled.

Why this matters

The sentence "the process won't die" actually lumps together at least four different situations.

  1. You sent kill but nothing changed -> the application caught SIGTERM and ignored it, or it is in the middle of cleanup
  2. Even kill -9 doesn't work -> the process is in the D state (uninterruptible wait). The signal is queued and is handled only once the process returns from the kernel
  3. It is already dead but still shows in the list -> it is a zombie. You cannot kill it, because it is already dead
  4. You stopped the container and data was cut off -> the signal stopped at the shell that is PID 1 and never reached the app

The response is different in all four situations. So do not stop at "it won't die"; ask at which layer it stopped.

How it works

A normal shutdown is always two steps. Send SIGTERM to give it time to clean up, and if it is still alive after the grace period, send SIGKILL. The default grace period of docker stop is 10 seconds, and systemd's TimeoutStopSec plays the same role. In a shell, the one line timeout -k 10 30 명령 (the last word stands for your own command) creates the same policy.

SIGKILL is dangerous because it gives no chance to clean up at all. Buffers are not flushed, transactions are cut off midway, and lock files and temporary files are left behind. For a database or a queue consumer, SIGKILL is effectively the same as pulling the plug.

It also helps to know the conventions for commonly used signals. Numbers can differ by architecture, so in scripts, always use the names.

Name Default action Conventional use
SIGHUP Terminate Reload configuration (originally meant the terminal connection dropped)
SIGINT Terminate Ctrl-C
SIGTERM Terminate A polite request to shut down
SIGKILL Terminate Cannot be caught. Last resort
SIGSTOP / SIGCONT Stop / Resume Pause for a moment and continue
SIGUSR1 / SIGUSR2 Terminate Up to the app to define (nginx: reopen logs / swap the binary)

Zombies and orphans are understood as a pair. When a child dies, the kernel keeps its exit status, and the parent must collect it with wait before it disappears from the list. The state before collection is a zombie (Z). A zombie uses almost no memory, but it occupies a PID, so when thousands pile up you reach the point where new processes cannot be created. Conversely, if the parent dies first, the child becomes an orphan and is adopted by PID 1, which collects it on the parent's behalf. This is why the sure way to get rid of zombies is "terminate the parent."

Remember the exit code convention as well. When a process dies from a signal, its exit code is 128 + 시그널 번호 (128 plus the signal number). SIGTERM is 143 and SIGINT is 130. If you see 143 in a CI log, read it as "someone politely asked it to shut down."

What you see in the field

The shell-form trap in containers. If you write CMD myapp in a Dockerfile, /bin/sh -c becomes PID 1 and the app becomes its child. The SIGTERM sent by the runtime is received by the shell, and the shell does not forward it to the child. After the grace period, SIGKILL is applied to the whole cgroup, and the app dies without doing any cleanup. The fix is the exec form (CMD ["myapp"]) or adding an init process.

The reach of pkill. pkill -f pattern-matches against the entire command line, so using it broadly will kill unrelated processes too. Always run pgrep -a -f with the same pattern first and check the list with your own eyes before running it. And never run kill -9 -1; it sends SIGKILL to every process your privileges reach.

Why trap runs late. In a shell script, a signal that arrives while sleep 300 is running is handled only after that command finishes. To make it react immediately, write it in the form sleep 300 & wait $!.

Processes that won't die and shutdowns that won't finish

The situation where you sent kill and the process does not die splits into three cases. Once you know which one it is, the next move is decided.

It is ignoring the signal. The program has a SIGTERM handler and is stuck doing cleanup inside it, or it was built to ignore the signal entirely. SIGKILL (9) cannot be caught by the process, so it always dies, but it skips cleanup — it does not close open files, does not release locks, and abandons writes in progress. If you use 9 first on a database or a queue consumer, recovery becomes necessary.

cat /proc/<pid>/status | grep -E 'SigBlk|SigIgn|SigCgt'

It is asleep inside the kernel in the D state. If the state in ps is D (uninterruptible sleep), it is waiting for a disk or NFS response, and it does not die even if you send 9. The kernel handles the signal only after it returns from that system call. What you need to fix here is not the process but the storage underneath it.

ps -eo pid,stat,wchan:20,cmd | awk '$2 ~ /D/'

It is already dead but the parent has not collected it. If the state is Z (zombie), the process has already finished and only its exit code remains. A zombie itself uses almost no resources, but when they pile up, PIDs run out. The one to kill is not the zombie but the parent. Zombies pile up in containers because PID 1 does not collect them, and --init or tini fills that role.

It decides the shutdown order by itself. A service that receives SIGTERM should stop accepting new requests, finish what is in progress, and only then exit. Without this order, requests are cut off at every deployment. If it still does not finish after some time, that is when killing it forcibly is right.

timeout -s TERM -k 30s 5m ./worker    # TERM 뒤 30초 안에 안 끝나면 KILL

Signals go to process groups. In a shell, Ctrl+C sends SIGINT to the entire foreground process group. The problem of a script dying and leaving its children behind is usually because the children are in a different group. Send to the whole group with kill -- -<pgid>.

What you will do in the next lab

You start a background process and check its PID and parent PID, then send stop, terminate, and force-kill signals to see how the state changes for each. You write a script that catches SIGTERM to prove for yourself that SIGKILL cannot be caught, create a zombie, and finally write a script that safely cleans up processes matching a pattern.