TT Lab
Get started
Learn Learning paths Courses

Docker Fundamentals

How a Container Dies

Continue in TT Lab

One-line summary

docker stop sends SIGTERM, waits 10 seconds by default, and then sends SIGKILL. Exit code 143 is SIGTERM (128+15) and 137 is SIGKILL (128+9). If you can tell these two numbers apart, you have already narrowed down half of the causes of an outage.

Why this is needed

There was a service whose shutdown took 10 seconds on every deployment and whose exit code was logged as 137. If you change only the command form of the same app and measure again, the result looks like this.

CMD npm start            ->  real 0m10.412s,  ExitCode 137
CMD ["node","server.js"] ->  real 0m0.284s,   ExitCode 143

10 seconds is Docker's default grace period, and 137 is a forced kill. With a clean shutdown it should have ended immediately with 143.

How it works

There are two layers to the cause.

First, the kernel treats PID 1 specially. For a signal without an explicitly registered handler, the default action is not applied. A shell-form CMD is actually run as /bin/sh -c "npm start", so the shell becomes PID 1, and since the shell has no SIGTERM handler, the signal is simply ignored.

Second, the explanation "the shell does not forward signals" is only half right. dash and busybox ash have an optimization: if sh -c contains only one simple command, they replace themselves with exec instead of forking. So the shell can disappear and the target can become PID 1. The catch is that this optimization is conditional. If even one pipe, &&, variable expansion or redirection gets in, the shell stays. Behavior that depends on the shell will silently break someday.

The right answer is an exec-form JSON array, and if you use an entrypoint script, the last line must be exec "$@".

What it looks like in the field

PID 1 has a second responsibility. A process that loses its parent is adopted by PID 1, and PID 1 must reap its exit status. An init system does this, but a Node or Python process does not. In a real case where the zombies were counted, 1847 of them had piled up, and once the process table is full, fork: Resource temporarily unavailable shows up in the application log.

Also, 137 may have been produced by the OOM killer. What separates the two is the State.OOMKilled field and the kernel log. An application cannot receive any exception for SIGKILL, so if a container keeps restarting silently, you have to look at the kernel log, not the application log.

Two cases where the signal never reaches PID 1

When you stop a container, the runtime sends SIGTERM to PID 1, and if it is still alive after --stop-timeout (10 seconds by default), it sends SIGKILL. But there are two cases where the signal does not reach the application.

Shell-form CMD. If you write CMD npm start, then /bin/sh -c "npm start" becomes PID 1, and the shell does not forward signals to its child.

CMD npm start                    # ❌ sh 가 PID 1. SIGTERM 을 삼킨다
CMD ["npm", "start"]             # ✅ exec 형식 — 애플리케이션이 PID 1

Wrapper script. If an entrypoint script calls the program at the end without exec, the script stays as PID 1.

#!/bin/sh
설정_준비
exec "$@"        # ← exec 이 있어야 프로세스가 교체되어 PID 1 이 된다

The symptom is the same. It does not react to the stop request and is forcibly killed 10 seconds later. Nothing is left in the log, and the requests and buffers that were in flight are lost.

Another duty of PID 1

On Linux, PID 1 also has the role of reaping orphan processes. If the application does not do this, zombie processes pile up. This is especially true for containers that create child processes (shell scripts, build tools).

docker run --init ...        # tini 를 PID 1 로 넣어 준다

In Kubernetes, if you use shareProcessNamespace: true, the pause container takes on that role; otherwise put tini in the image. If ps shows Z states piling up, this is the problem.

Designing a graceful shutdown

1. SIGTERM 수신
2. 새 요청 받기를 멈춘다 (헬스 체크를 실패로 바꾼다)
3. 진행 중 요청이 끝나기를 기다린다
4. 연결·파일을 닫고 종료

Step 2 is especially important. When Kubernetes deletes a Pod, it sends the endpoint removal and SIGTERM at the same time, but it takes a few seconds for the endpoint removal to propagate. Requests that arrive in that window go to the Pod that is shutting down and fail. In the preStop hook, adding a short sleep closes that window.

lifecycle:
  preStop:
    exec: {command: ["sh", "-c", "sleep 5"]}
terminationGracePeriodSeconds: 30

What you will do in the next lab

You will start one container with a handler and one without, produce 143 and 137 yourself, and also check the application exit code and the restart policy.