Kubernetes Distributions — Build Them Yourself
I ran the snapshot command and the datastore was SQLite
Goal
You confirm that the datastore of a single k3s server is SQLite (kine), see why an etcd snapshot is rejected, and then move to the embedded etcd with a --cluster-init restart. You back up a snapshot together with the token and restore it with --cluster-reset, and record that the changes made after the snapshot disappear and where the previous data remains.
Why it matters
When k3s starts with one machine, it uses SQLite instead of etcd. It is light and fast, but several servers cannot share it, and etcd tools such as k3s etcd-snapshot do not work.
So it is on the day you say "let's set up backups" that you check for the first time what the datastore is, and your options then are an SQLite backup that copies the files wholesale, or moving to etcd.
Moving to etcd takes one restart, but a restore is "rolling back to that point in time", so what you created after the snapshot disappears. And without the server token, the snapshot is useless.
This lab has you experience those three things — checking the datastore, the migration, and the scope of what a restore reverts — once each on a real k3s. Unlike the k0s lab, which uses etcd from the start, it deals with the process of changing the datastore itself.
Steps
- In the namespace
k3s-ds, create a ConfigMapmarker(stage=sqlite). Then write to/root/k3s-ds/datastore.jsonthe fieldsdatastore(sqliteoretcd),db_files(a sorted array of the file names in/var/lib/rancher/k3s/server/db),kine_table(among the table names in state.db, the one kine uses),kine_endpoint(the address afterKine available atin the k3s log), andmarker_uid(the UID of marker). - In the current state, run
k3s etcd-snapshot saveand save the standard output and standard error together to/root/k3s-ds/snapshot-refused.txt. Then write to/root/k3s-ds/sqlite-backup.txt, one per line, the absolute paths of the directory and file you would have to copy to back up the SQLite datastore. - Write
cluster-init: truein/etc/rancher/k3s/config.yamland restart k3s. Confirm that marker is still alive and write to/root/k3s-ds/migrate.jsonthe fieldsmarker_uid(the UID of marker after the restart),sqlite_file_now(the name the original state.db was changed to in the db directory),node_roles(a sorted array of the node-role.kubernetes.io/ label names of the node, for example ["control-plane"]), andmember_name(the content of thedb/etcd/namefile). - Change the
stageof marker tosnapshotted, and then take a snapshot withk3s etcd-snapshot save --name lab-before. Write to/root/k3s-ds/snapshot.jsonthe fieldsname(the full name of the snapshot that was created),path(the absolute path of the file),size(in bytes, an integer), andsha256(the SHA-256 of the file). - Copy the snapshot file from step 4 (with the same name) and the server token file (named
token, permission 600) into the/root/k3s-ds/backup/directory. Write to/root/k3s-ds/token-backup.jsonthe fieldstoken_source(the absolute path of the token original),token_sha256(the SHA-256 of the token file), andsnapshot_sha256(the SHA-256 of the backed-up snapshot). Do not write the token value itself anywhere. - After the snapshot, create a ConfigMap
after-snap(x=1) in the namespacek3s-dsand change thestageof marker tochanged. Write to/root/k3s-ds/after.jsonthe fieldsafter_uid(the UID of after-snap),after_created(its creationTimestamp), andmarker_stage(the current stage of marker). - Stop k3s, restore with
k3s server --cluster-reset --cluster-reset-restore-path=<4단계 스냅숏 경로>(the placeholder is the path of the snapshot from step 4) (save the entire output to/root/k3s-ds/reset.log), and then start k3s again. Write to/root/k3s-ds/restore.jsonthe fieldsold_dir(the name of the directory to which the restore moved the previous etcd data),restart_hint(from the sentences in reset.log in which k3s tells you what to do next, the first sentence, up to the part that containsrestart without),after_snap_exists(a boolean),marker_stage(the stage of marker after the restore), andmember_name(the content ofdb/etcd/nameafter the restore). - Write to
/root/k3s-ds/report.jsonthe fieldsdatastore_now(sqliteoretcd),lost_objects(a sorted array of the ConfigMap names that disappeared because of the restore),marker_stage(the current value),snapshot_dir(the absolute path of the directory where the snapshot is stored),restore_needs_token(whether the original token is needed when restoring on a different server, a boolean), andmember_name_changed(whether the etcd member name changed between before and after the restore, a boolean).
Notes
- There is one k3s v1.35.8+k3s1 in the VM (traefik and metrics-server are off). A restart is
systemctl restart k3s. - After step 3, the
k3s etcd-snapshotcommand prints one warning line saying cluster-init in config.yaml is an unknown key. It is harmless (measured). - As in the documented procedure, run the restore command after stopping the service with
systemctl stop k3s. - Common mistake: backing up only the snapshot and leaving out
/var/lib/rancher/k3s/server/token. - Common mistake: putting
--cluster-resetin the service arguments or config.yaml. A restore is a command you run separately once, and k3s prevents consecutive resets with the reset-flag file. - Cluster Datastore · High Availability Embedded etcd · etcd-snapshot · Backup and Restore
What is the datastore of this cluster
In the namespace k3s-ds, create a ConfigMap marker (stage=sqlite). Then write to /root/k3s-ds/datastore.json the fields datastore (sqlite or etcd), db_files (a sorted array of the file names in /var/lib/rancher/k3s/server/db), kine_table (among the table names in state.db, the one kine uses), kine_endpoint (the address after Kine available at in the k3s log), and marker_uid (the UID of marker).
k3s does not attach SQLite directly to the API server; it puts kine, which imitates the etcd API, in between. Check in the log what the API server's --etcd-servers points to. You can see the table names with sqlite3 <파일> .tables (the placeholder is the database file).
You tried the snapshot command but the datastore was SQLite
In the current state, run k3s etcd-snapshot save and save the standard output and standard error together to /root/k3s-ds/snapshot-refused.txt. Then write to /root/k3s-ds/sqlite-backup.txt, one per line, the absolute paths of the directory and file you would have to copy to back up the SQLite datastore.
etcd-snapshot sends a request to the server, and the server takes the snapshot with the embedded etcd. If the datastore is SQLite, the server rejects it and leaves the detailed reason in the server log. SQLite is backed up by copying files without any special command, but the value that encrypts the confidential data in the datastore must be kept along with it.
Move to etcd with a single restart
Write cluster-init: true in /etc/rancher/k3s/config.yaml and restart k3s. Confirm that marker is still alive and write to /root/k3s-ds/migrate.json the fields marker_uid (the UID of marker after the restart), sqlite_file_now (the name the original state.db was changed to in the db directory), node_roles (a sorted array of the node-role.kubernetes.io/ label names of the node, for example ["control-plane"]), and member_name (the content of the db/etcd/name file).
If you start a server that was running on SQLite with cluster-init, k3s moves the SQLite content to etcd. If etcd data is already on disk, this argument is ignored. Look for Migrating content from sqlite to etcd in the log.
Take a named snapshot
Change the stage of marker to snapshotted, and then take a snapshot with k3s etcd-snapshot save --name lab-before. Write to /root/k3s-ds/snapshot.json the fields name (the full name of the snapshot that was created), path (the absolute path of the file), size (in bytes, an integer), and sha256 (the SHA-256 of the file).
--name decides only the first part of the name, and k3s appends the node name and the time. The storage location is the default of --etcd-snapshot-dir. k3s etcd-snapshot ls and kubectl get etcdsnapshotfile show the same snapshots.
You cannot restore with the snapshot alone
Copy the snapshot file from step 4 (with the same name) and the server token file (named token, permission 600) into the /root/k3s-ds/backup/ directory. Write to /root/k3s-ds/token-backup.json the fields token_source (the absolute path of the token original), token_sha256 (the SHA-256 of the token file), and snapshot_sha256 (the SHA-256 of the backed-up snapshot). Do not write the token value itself anywhere.
k3s encrypts the confidential bootstrap data in the datastore with the server token. If you restore with a different token, the snapshot cannot be used. Make sure the permissions do not get wider when you copy.
What gets created after the snapshot
After the snapshot, create a ConfigMap after-snap (x=1) in the namespace k3s-ds and change the stage of marker to changed. Write to /root/k3s-ds/after.json the fields after_uid (the UID of after-snap), after_created (its creationTimestamp), and marker_stage (the current stage of marker).
The point is to see what the restore in the next step does to these two. Write the UID and time exactly so that you can confirm after the restore that this record was genuine.
After the restore, the changes made after the snapshot disappeared
Stop k3s, restore with k3s server --cluster-reset --cluster-reset-restore-path=<4단계 스냅숏 경로> (the placeholder is the path of the snapshot from step 4) (save the entire output to /root/k3s-ds/reset.log), and then start k3s again. Write to /root/k3s-ds/restore.json the fields old_dir (the name of the directory to which the restore moved the previous etcd data), restart_hint (from the sentences in reset.log in which k3s tells you what to do next, the first sentence, up to the part that contains restart without), after_snap_exists (a boolean), marker_stage (the stage of marker after the restore), and member_name (the content of db/etcd/name after the restore).
A restore means running the same binary once separately while the service is stopped. When it finishes it exits by itself and tells you to start again. It does not delete the previous data but moves it aside. A marker file that prevents consecutive resets is created and then deleted after a normal startup.
What remained and what disappeared
Write to /root/k3s-ds/report.json the fields datastore_now (sqlite or etcd), lost_objects (a sorted array of the ConfigMap names that disappeared because of the restore), marker_stage (the current value), snapshot_dir (the absolute path of the directory where the snapshot is stored), restore_needs_token (whether the original token is needed when restoring on a different server, a boolean), and member_name_changed (whether the etcd member name changed between before and after the restore, a boolean).
Write them based on the json files from earlier steps and the current disk and cluster. The grader recalculates the same values from the record files, the etcd-old directory, and the cluster.