Data Must Outlive the Container
One-line summary
A container's writable layer dies with the container. Data that must live long should be kept outside the container, in a volume or a bind mount.
Why this is needed
Reports that "all the uploaded files disappeared after I redeployed the container" are still common today. The cause is not a bug but the design. The writable layer is attached to the container object, and when you delete the container, that layer disappears with it.
A performance problem is layered on top of this. To modify a file in a lower layer, the whole file has to be copied up into the upper layer (copy-up), which is very bad for large files such as database files. The saying "always move large writes to a volume" has become an idiom for this reason.
How it works
There are three methods, each with its own character.
| Item | Bind mount | Named volume | tmpfs |
|---|---|---|---|
| Location | You specify a host path directly | Managed by the runtime | Memory |
| Persistence | Permanent on the host | Permanent in the volume | Gone on exit |
| Portability | Low (depends on the path) | High | High |
| Backup | Managed by hand | Dedicated command | Not possible |
| Main use | Configuration files, source during development | DB data, uploads | Temporary files, secrets |
There are also two syntaxes: the short -v source:target:opts and the explicit --mount type=...,source=...,target=...,readonly. The latter exposes the option names, so there is less room for misunderstanding in reviews.
A read-only mount matters more than you might think. If you attach a configuration file with :ro, the container cannot modify the configuration, whether by mistake or through a compromise. If the application tries to write, a Read-only file system error occurs, and this error is itself a signal that "this container is trying to write somewhere it must not".
What it looks like in the field
The backup pattern has hardened into an idiom. You start a temporary container that mounts both the volume and the backup directory, and create an archive.
docker run --rm -v dk-data:/source:ro -v /root/backup:/backup alpine:3.20 tar czf /backup/dk-data.tgz -C /source .
The key is to use -C to enter the source directory and store files with relative paths. Only then do the paths not get misaligned during restore.
In rootless environments, ownership problems come up often. This is because the UID mapping differs between the host and the container, and in that case, instead of opening the files with chmod 777, it is right to understand the mapping and match it. If you leave the permissions open, you are opening them not only to that container but also to other processes on the same host.
Criteria for choosing between a volume and a bind mount
The names are similar, but the characteristics differ.
| Named volume | bind mount | |
|---|---|---|
| Managed by | Docker | A person |
| Location | /var/lib/docker/volumes/… |
Any path on the host |
| Backup | Extract with docker run --rm -v vol:/d … |
The host files as they are |
| Permissions | Initialized following the image's UID | The host permissions as they are |
| Development convenience | Source edits are not visible | Reflected immediately |
The permissions row is the practical trap. A bind mount takes the owner and mode of the host file as they are, so if the container runs as nonroot (UID 65532) and the host file is root:root 0644, it cannot write. A named volume, on the other hand, is initialized when first created by copying the permissions of the target path inside the container, so it usually just works.
It is copied only the first time
A named volume is initialized only once, when it is empty.
빈 볼륨을 /app/config 에 붙이면
→ 이미지의 /app/config 내용이 볼륨으로 복사된다
이미 내용이 있는 볼륨을 붙이면
→ 이미지의 내용은 가려진다. 복사하지 않는다
So even if you build a new image and change the default configuration file, a container that uses the old volume keeps seeing the old file. This is a common cause of "I fixed the image but it is not reflected". It is better to inject configuration through ConfigMaps or environment variables rather than keeping it in a volume.
The Kubernetes equivalents
| Docker | Kubernetes | Where it is used |
|---|---|---|
| Named volume | PVC | DB data |
| bind mount | hostPath | Node file access (avoid if possible) |
| tmpfs | emptyDir: {medium: Memory} |
Secrets, temporary data |
| — | emptyDir: {} |
Sharing between containers (Pod lifetime) |
emptyDir disappears when the Pod is deleted. It also disappears when the node reboots. But it survives a Pod restart (when only the container dies), which makes it easy to be fooled when testing.
Solve permission problems with fsGroup.
spec:
securityContext:
fsGroup: 2000 # 볼륨의 그룹 소유자를 2000 으로 바꿔 준다
runAsUser: 1000
runAsNonRoot: true
fsGroup recursively changes the ownership of every file in the volume, so if there are very many files, Pod startup becomes slow. In that case, use fsGroupChangePolicy: OnRootMismatch so that only the root directory is checked.
What you will do in the next lab
You will create and use a volume, check that the data remains even after you delete the container, produce the error message of a read-only mount yourself, and then do a round trip of backup and restore.