What Do You Look At First When a Container Dies
One-line summary
Diagnosing a container means looking at three things in order: logs, then the exit code, then the state fields. If these three do not solve it, only then do you go inside the container.
Why this is needed
The report "the container won't start" actually lumps together at least four different situations into one sentence: the image is missing, the command is missing, the app could not read its configuration and ended itself, or the kernel killed it. All four call for completely different responses, and fortunately they can be told apart with very cheap signals.
The exit code narrows the scope first
One exit code halves the scope of the investigation. The rule is simple. If it is greater than 128, the process died from a signal, and subtracting 128 gives the signal number.
| Code | Meaning | Where to look first |
|---|---|---|
| 0 | Normal exit | The app finished its work and ended. For a server, even this is strange |
| 1 | General error from the app | The last line of the log |
| 125 | Error in Docker itself | The command options are wrong |
| 126 | Command cannot be executed | No execute permission |
| 127 | Command not found | A typo in the ENTRYPOINT path, or an image without a shell |
| 137 | SIGKILL (128+9) | You must tell whether it was OOM or not |
| 139 | SIGSEGV (128+11) | A native library crash |
| 143 | SIGTERM (128+15) | It received a request for a clean shutdown. It may not be an incident |
137 is the most confusing. It may have been killed by the kernel's OOM killer, or a person may have run docker kill. The way to tell is the single line State.OOMKilled.
docker inspect dk-err --format '{{.State.ExitCode}} {{.State.OOMKilled}}'
137 true
If it is true, the problem is raising the memory limit or reducing the app's usage; if it is false, the problem is finding out who killed it and why. These are completely different investigations.
How to read logs
The runtime intercepts and stores a container's standard output and standard error. So even if the container is already dead, you can see its last moments with docker logs. The two streams appear mixed together, but they are actually separate, so you can pull them out separately with shell redirection.
docker logs dk-err 2>/dev/null # 표준 출력만
docker logs dk-err 2>&1 1>/dev/null # 표준 오류만
docker logs --tail 50 -t dk-err # 마지막 50줄에 시각을 붙여서
docker logs --since 10m dk-err # 최근 10분만
Adding timestamps (-t) is important. Many apps do not put timestamps in their logs, and then you cannot tell whether an error is from right before the crash or from long before.
docker inspect holds the entire state in one JSON document. Only a few fields are actually used in diagnosis.
| Field | What it tells you |
|---|---|
State.Status |
created / running / exited / paused |
State.ExitCode |
Whether the app produced the value or it died from a signal |
State.Error |
Why the runtime could not start it at all |
State.OOMKilled |
Whether 137 was caused by OOM |
State.StartedAt / FinishedAt |
How many seconds it lived — if it died instantly, it is a configuration problem |
RestartCount |
Whether it is silently restarting over and over |
If the difference between StartedAt and FinishedAt is less than 1 second, the problem is not the app logic but the startup conditions. Check in this order: configuration files, environment variables, port conflicts.
What it looks like in the field
The trap people step into most often is an empty log. There are three causes.
- The app writes its logs to a file. The principle for containers is to send logs to standard output, not to files, and if this rule is not followed, the whole diagnostic toolchain becomes useless. As a stopgap you can
docker execin and read the file, but once the container dies you cannot even do that. - It is stuck in buffering. Python uses block buffering when standard output is not a terminal. If it dies before 4 KB fills up, the logs inside are lost entirely. Turn it off with
PYTHONUNBUFFERED=1orpython -u. Node is unbuffered by default, and Java is usually fine becauseSystem.outis line-buffered. - The log driver is different. If
--log-driverisnoneor is set to an external collector,docker logscannot show anything.
The second trap is an image without a shell. A distroless or scratch-based image has no sh, so docker exec -it ... sh does not work. If you see OCI runtime exec failed: exec: "sh": executable file not found and conclude that "something is wrong with the container", you are mistaken. This is normal.
In this case you attach a container that has the tools into the same namespaces and look inside.
docker run -it --rm --pid container:myapp --network container:myapp nicolaka/netshoot
--pid shares the processes and --network shares the network, so ps and ss show those of the target container as they are. The filesystem is not shared, so you look at it through /proc/1/root/.
The third is repeated restarts. If RestartCount keeps growing, docker logs shows only the logs of the current instance. You need the logs of the instance that just died, but Docker has no option like Kubernetes --previous. Temporarily change the --restart policy to no so that it starts only once, and investigate.
What you will do in the next lab
You will cut a log from the tail, separate the two streams, and practice identifying the cause by putting the exit code and log of a dead container side by side. You will also set a low memory limit to produce 137 yourself and confirm that OOMKilled shows true.