Kubernetes Distributions — Build Them Yourself
k3s starts with SQLite as its datastore
One-line summary
A single k3s server uses SQLite through kine by default, so etcd snapshots do not work; restarting it with --cluster-init moves it to the embedded etcd while preserving the data; and restoring an etcd snapshot requires the server token and reverts every change made after the snapshot.
Why this was needed
The Kubernetes API server was built to store its state in etcd. But etcd is a distributed store designed on the assumption of several machines for quorum, so it is heavy for a single Raspberry Pi or a cluster used briefly in CI.
k3s bridges this gap with kine. kine is a thin layer that imitates the etcd API and puts an ordinary database such as SQLite, MySQL, or PostgreSQL behind it. If you look at the logs in measurement, the API server is running with --etcd-servers=unix://kine.sock, and k3s announces Kine available at unix://kine.sock first.
The official documentation says that SQLite is the default used when no other datastore is configured and no embedded etcd data exists on disk, and that it cannot be used in a cluster with several servers.
The problem is that this default is quiet. The cluster runs fine, and it only shows up on the day you try to set up backups.
$ k3s etcd-snapshot save
level=fatal msg="Error: see server log for details: etcd datastore disabled"
This is the measured output. The same message is left in the server log as an HTTP 400 response. The snapshot command sent a request to the k3s server, and the server rejected it, saying "I do not use etcd."
How it works
Backup with SQLite has no special command. As the documentation says, you copy /var/lib/rancher/k3s/server/db/, and when restoring, you put that content back. Along with it, you must also keep the server token file /var/lib/rancher/k3s/server/token. The token is used to encrypt the confidential data in the datastore, so if you restore with a different token, the backup cannot be used.
Moving to etcd is nothing more than starting a server that was running on SQLite again with --cluster-init. The measured result is as follows.
config.yaml 에 cluster-init: true → systemctl restart k3s (API 준비 12초)
로그 Migrating content from sqlite to etcd
디스크 db/state.db → db/state.db.migrated, db/etcd/ 와 db/snapshots/ 생성
노드 node-role.kubernetes.io/etcd=true 가 붙음
데이터 재시작 전에 만든 ConfigMap 의 UID 가 그대로
There is no reverse direction. The documentation says that if etcd data is found on disk, datastore arguments such as --cluster-init, --server, and --datastore-endpoint are ignored. This means that once it has become etcd, it does not go back to SQLite even if you remove the argument.
Snapshots come in two kinds, scheduled and manual. Scheduled snapshots are taken by default at 00:00 and 12:00 (0 */12 * * *), 5 are retained, and the name is etcd-snapshot-<노드>-<시각> (that is, the node name and the time). Manual ones are taken with k3s etcd-snapshot save, have no limit on the number kept so you must delete them yourself, and --name decides only the first part of the name. Both are saved under the data directory, the default of --etcd-snapshot-dir, in db/snapshots, and the measured path was /var/lib/rancher/k3s/server/db/snapshots/lab-before-<노드>-<유닉스시각> (node name and Unix time). kubectl get etcdsnapshotfile shows the snapshots of the whole cluster as objects.
Restore is done by stopping the service and running the same binary once separately.
systemctl stop k3s
k3s server --cluster-reset --cluster-reset-restore-path=/var/lib/rancher/k3s/server/db/snapshots/<이름>
# Managed etcd cluster membership has been reset, restart without --cluster-reset flag now.
systemctl start k3s
The restore does not delete the current etcd data; it moves it to db/etcd-old-<시각> (the placeholder is the time), then unpacks the snapshot, removes all the other members, and makes it a cluster of one. To prevent consecutive resets it creates db/reset-flag and deletes it on a normal startup. If you give only --cluster-reset without --cluster-reset-restore-path, only the membership is reset without any snapshot restore.
What it looks like in the field
In measurement, after taking a snapshot, I created one ConfigMap, changed the value of another, and then restored. The newly created one disappeared and the value went back to the snapshot point. The UID of the vanished object remained as bytes in the old data file inside etcd-old-<시각> (the placeholder is the time) — this is how it is confirmed that a restore does not delete the old data. The etcd member name (db/etcd/name) was also newly assigned, a different value from before the restore.
There are two common incidents in the field. One is "I restored and what we deployed yesterday is gone." A restore is a rewind, so if you don't first check who did what after the snapshot, it becomes a second outage. The other is failing when you move only the snapshot to a new server and try to restore. The documentation says that when restoring on a different host, you must pass the original server's token with --token. If the snapshot is in object storage but the token was only on the vanished node's disk, that backup is unusable.
What really matters in practice
- When you take over a cluster, first check what the datastore is. You only need to see whether
db/hasstate.dboretcd/. - If you plan to add servers, start with cluster-init from the beginning. The documentation explains that the embedded etcd must consist of an odd number of servers for quorum.
- A backup is snapshot plus token as one bundle. Keep the token with permission 600 and also in a different place from the snapshot.
- Before restoring, record "what changed after the snapshot." It becomes the list of things to apply again after the restore.
- A restore leaves the old data as
etcd-old-<시각>(the placeholder is the time). It takes up disk space, so clean it up after the restore is confirmed. - According to the documentation, when restoring, the k3s version need not be the same as the one that created the snapshot, and a higher minor version is also accepted.
What you will do in the next lab
In k3s inside the VM, you confirm from the kine log and the table that the datastore is SQLite, and record that the snapshot is rejected. You move to etcd with a cluster-init restart and check whether the UID of the marker ConfigMap survives, back up a named snapshot together with the token, then make changes after the snapshot and restore, and check what disappears and where it remains.
Reference documents: Cluster Datastore · High Availability Embedded etcd · etcd-snapshot · Backup and Restore