TT Lab
Get started
Learn Learning paths Courses

Volumes, Networks and Compose

The Container Is Alive and the Service Is Dead

Continue in TT Lab

One-line summary

A process being alive and its being able to handle requests are different facts. A health check fills the gap between the two, and a restart policy decides what to do on failure.

Why this is needed

docker ps says Up 3 hours. But users are getting 500s. The process is alive, but the DB connection pool is exhausted, or a thread is deadlocked, or the disk is full and it cannot write.

The only thing Docker knows is one fact: PID 1 has not died yet. If you judge "in service" from that alone, you miss outages for hours.

How it works

HEALTHCHECK --interval=30s --timeout=3s --start-period=40s --retries=3 \
  CMD curl -fsS http://localhost:8080/healthz || exit 1
Option Meaning If set wrongly
interval Check period Too short causes load, too long delays detection
timeout Time limit for one check Too short judges a healthy service as failed
start-period Startup grace — failures in this window do not count as retries Without it, an app that starts slowly keeps dying
retries Number of consecutive failures allowed With 1, even a momentary delay makes it unhealthy

The state moves from starting to healthy / unhealthy. You can see it with docker inspect --format '{{.State.Health.Status}}'.

An important trap: with Docker alone, even when a container becomes unhealthy, it does not restart the container. It only displays the state. Restarting is done by an orchestrator (Compose's depends_on: condition: service_healthy, Swarm, Kubernetes' liveness probe). Do not believe that "we added a health check, so it will revive by itself".

How to use a health check endpoint

If /healthz unconditionally returns just 200 OK, it is the same as checking nothing. Conversely, if you check the DB, cache and external APIs all together, then when one external API wobbles, our whole service becomes unhealthy.

Draw the boundary like this.

If you do not make this distinction, you fall into the feedback loop "the DB got slow for a moment → liveness fails → restart → it gets even slower reconnecting → fails again". This really happens often.

Restart policies

Policy Behavior
no (default) Leaves it as is when it dies
on-failure[:N] Only on a non-zero exit code, at most N times
always Always restarts. Also comes back up when the daemon restarts
unless-stopped Same as always, but leaves alone what a person stopped

always can create an infinite restart loop. The typical case is a container that dies immediately because of a configuration error, restarting several times a second and filling the log. Docker mitigates this by gradually lengthening the restart interval (backoff), but it is not a fundamental fix.

depends_on guarantees only order

services:
  app:
    depends_on: [db]     # db 컨테이너가 '시작'된 뒤 app 을 시작할 뿐

It does not look at whether db is ready to accept requests. PostgreSQL takes a few more seconds to initialize even after the process is up. If the app connects in between, it fails. It only gains meaning when you tie it to the health check with condition: service_healthy.

When a health check causes an outage instead

A health check is a safety mechanism, but if made wrongly it is worse than having none. Avoid these three.

If you put dependencies in the check, they die together. If /health queries the database, when the DB slows down for a moment, every instance becomes unhealthy at the same time and restarts. The restarted instances attach to the DB all at once, and the situation gets worse. A check that asks whether it is alive looks at itself only.

It kills a program that starts slowly. A service that loads a JVM or a large model takes 1 minute to come up, and if the check fails at 30 seconds, it only keeps restarting forever. Docker gives this time with --start-period, and Kubernetes with startupProbe.

HEALTHCHECK --interval=30s --timeout=3s --start-period=60s --retries=3   CMD curl -fsS http://localhost:8080/healthz || exit 1

The restart policy hides the failure. restart: always revives it every time it dies, so a service that dies every 30 seconds because of a wrong configuration looks, on the surface, like it is "running". If you set a count like on-failure:5, it stops when it keeps failing, and then a person notices. Kubernetes' CrashLoopBackOff means the same thing — it is a mechanism that slows down exponentially while exposing the problem.

The tool used for the check may not be in the image. A distroless image has neither curl nor wget, so the health check always fails. In that case, either give the application binary a mode in which it checks itself, or use a method that checks from the outside, like Kubernetes' HTTP probe.

Separate the two checks. Whether it is ready to receive traffic (readiness) and whether it is alive (liveness) are different questions. If the former fails, only the traffic is pulled away, and if the latter fails, it restarts. If you merge them into one, an instance that is briefly busy gets restarted.

What it looks like in the field

What to look for in the next check

In the quiz that follows, you will distinguish running (running) from healthy (healthy), and judge which failures to apply start-period, restart policies and Compose's health conditions to. Answer by citing the state transitions and exit codes of the containers you built in the previous lab.