TT Lab
Get started
Learn Learning paths Courses

Docker Fundamentals

What Do You Look At First When a Container Dies

Continue in TT Lab

One-line summary

Diagnosing a container means looking at three things in order: logs, then the exit code, then the state fields. If these three do not solve it, only then do you go inside the container.

Why this is needed

The report "the container won't start" actually lumps together at least four different situations into one sentence: the image is missing, the command is missing, the app could not read its configuration and ended itself, or the kernel killed it. All four call for completely different responses, and fortunately they can be told apart with very cheap signals.

The exit code narrows the scope first

One exit code halves the scope of the investigation. The rule is simple. If it is greater than 128, the process died from a signal, and subtracting 128 gives the signal number.

Code Meaning Where to look first
0 Normal exit The app finished its work and ended. For a server, even this is strange
1 General error from the app The last line of the log
125 Error in Docker itself The command options are wrong
126 Command cannot be executed No execute permission
127 Command not found A typo in the ENTRYPOINT path, or an image without a shell
137 SIGKILL (128+9) You must tell whether it was OOM or not
139 SIGSEGV (128+11) A native library crash
143 SIGTERM (128+15) It received a request for a clean shutdown. It may not be an incident

137 is the most confusing. It may have been killed by the kernel's OOM killer, or a person may have run docker kill. The way to tell is the single line State.OOMKilled.

docker inspect dk-err --format '{{.State.ExitCode}} {{.State.OOMKilled}}'
137 true

If it is true, the problem is raising the memory limit or reducing the app's usage; if it is false, the problem is finding out who killed it and why. These are completely different investigations.

How to read logs

The runtime intercepts and stores a container's standard output and standard error. So even if the container is already dead, you can see its last moments with docker logs. The two streams appear mixed together, but they are actually separate, so you can pull them out separately with shell redirection.

docker logs dk-err 2>/dev/null      # 표준 출력만
docker logs dk-err 2>&1 1>/dev/null # 표준 오류만
docker logs --tail 50 -t dk-err     # 마지막 50줄에 시각을 붙여서
docker logs --since 10m dk-err      # 최근 10분만

Adding timestamps (-t) is important. Many apps do not put timestamps in their logs, and then you cannot tell whether an error is from right before the crash or from long before.

docker inspect holds the entire state in one JSON document. Only a few fields are actually used in diagnosis.

Field What it tells you
State.Status created / running / exited / paused
State.ExitCode Whether the app produced the value or it died from a signal
State.Error Why the runtime could not start it at all
State.OOMKilled Whether 137 was caused by OOM
State.StartedAt / FinishedAt How many seconds it lived — if it died instantly, it is a configuration problem
RestartCount Whether it is silently restarting over and over

If the difference between StartedAt and FinishedAt is less than 1 second, the problem is not the app logic but the startup conditions. Check in this order: configuration files, environment variables, port conflicts.

What it looks like in the field

The trap people step into most often is an empty log. There are three causes.

  1. The app writes its logs to a file. The principle for containers is to send logs to standard output, not to files, and if this rule is not followed, the whole diagnostic toolchain becomes useless. As a stopgap you can docker exec in and read the file, but once the container dies you cannot even do that.
  2. It is stuck in buffering. Python uses block buffering when standard output is not a terminal. If it dies before 4 KB fills up, the logs inside are lost entirely. Turn it off with PYTHONUNBUFFERED=1 or python -u. Node is unbuffered by default, and Java is usually fine because System.out is line-buffered.
  3. The log driver is different. If --log-driver is none or is set to an external collector, docker logs cannot show anything.

The second trap is an image without a shell. A distroless or scratch-based image has no sh, so docker exec -it ... sh does not work. If you see OCI runtime exec failed: exec: "sh": executable file not found and conclude that "something is wrong with the container", you are mistaken. This is normal.

In this case you attach a container that has the tools into the same namespaces and look inside.

docker run -it --rm   --pid container:myapp --network container:myapp   nicolaka/netshoot

--pid shares the processes and --network shares the network, so ps and ss show those of the target container as they are. The filesystem is not shared, so you look at it through /proc/1/root/.

The third is repeated restarts. If RestartCount keeps growing, docker logs shows only the logs of the current instance. You need the logs of the instance that just died, but Docker has no option like Kubernetes --previous. Temporarily change the --restart policy to no so that it starts only once, and investigate.

What you will do in the next lab

You will cut a log from the tail, separate the two streams, and practice identifying the cause by putting the exit code and log of a dead container side by side. You will also set a low memory limit to produce 137 yourself and confirm that OOMKilled shows true.