TT Lab
Get started
Learn Learning paths Courses

Kubernetes Distributions — Build Them Yourself

I ran the snapshot command and the datastore was SQLite

Continue in TT Lab

Goal

You confirm that the datastore of a single k3s server is SQLite (kine), see why an etcd snapshot is rejected, and then move to the embedded etcd with a --cluster-init restart. You back up a snapshot together with the token and restore it with --cluster-reset, and record that the changes made after the snapshot disappear and where the previous data remains.

Why it matters

When k3s starts with one machine, it uses SQLite instead of etcd. It is light and fast, but several servers cannot share it, and etcd tools such as k3s etcd-snapshot do not work. So it is on the day you say "let's set up backups" that you check for the first time what the datastore is, and your options then are an SQLite backup that copies the files wholesale, or moving to etcd. Moving to etcd takes one restart, but a restore is "rolling back to that point in time", so what you created after the snapshot disappears. And without the server token, the snapshot is useless. This lab has you experience those three things — checking the datastore, the migration, and the scope of what a restore reverts — once each on a real k3s. Unlike the k0s lab, which uses etcd from the start, it deals with the process of changing the datastore itself.

Steps

  1. In the namespace k3s-ds, create a ConfigMap marker (stage=sqlite). Then write to /root/k3s-ds/datastore.json the fields datastore (sqlite or etcd), db_files (a sorted array of the file names in /var/lib/rancher/k3s/server/db), kine_table (among the table names in state.db, the one kine uses), kine_endpoint (the address after Kine available at in the k3s log), and marker_uid (the UID of marker).
  2. In the current state, run k3s etcd-snapshot save and save the standard output and standard error together to /root/k3s-ds/snapshot-refused.txt. Then write to /root/k3s-ds/sqlite-backup.txt, one per line, the absolute paths of the directory and file you would have to copy to back up the SQLite datastore.
  3. Write cluster-init: true in /etc/rancher/k3s/config.yaml and restart k3s. Confirm that marker is still alive and write to /root/k3s-ds/migrate.json the fields marker_uid (the UID of marker after the restart), sqlite_file_now (the name the original state.db was changed to in the db directory), node_roles (a sorted array of the node-role.kubernetes.io/ label names of the node, for example ["control-plane"]), and member_name (the content of the db/etcd/name file).
  4. Change the stage of marker to snapshotted, and then take a snapshot with k3s etcd-snapshot save --name lab-before. Write to /root/k3s-ds/snapshot.json the fields name (the full name of the snapshot that was created), path (the absolute path of the file), size (in bytes, an integer), and sha256 (the SHA-256 of the file).
  5. Copy the snapshot file from step 4 (with the same name) and the server token file (named token, permission 600) into the /root/k3s-ds/backup/ directory. Write to /root/k3s-ds/token-backup.json the fields token_source (the absolute path of the token original), token_sha256 (the SHA-256 of the token file), and snapshot_sha256 (the SHA-256 of the backed-up snapshot). Do not write the token value itself anywhere.
  6. After the snapshot, create a ConfigMap after-snap (x=1) in the namespace k3s-ds and change the stage of marker to changed. Write to /root/k3s-ds/after.json the fields after_uid (the UID of after-snap), after_created (its creationTimestamp), and marker_stage (the current stage of marker).
  7. Stop k3s, restore with k3s server --cluster-reset --cluster-reset-restore-path=<4단계 스냅숏 경로> (the placeholder is the path of the snapshot from step 4) (save the entire output to /root/k3s-ds/reset.log), and then start k3s again. Write to /root/k3s-ds/restore.json the fields old_dir (the name of the directory to which the restore moved the previous etcd data), restart_hint (from the sentences in reset.log in which k3s tells you what to do next, the first sentence, up to the part that contains restart without), after_snap_exists (a boolean), marker_stage (the stage of marker after the restore), and member_name (the content of db/etcd/name after the restore).
  8. Write to /root/k3s-ds/report.json the fields datastore_now (sqlite or etcd), lost_objects (a sorted array of the ConfigMap names that disappeared because of the restore), marker_stage (the current value), snapshot_dir (the absolute path of the directory where the snapshot is stored), restore_needs_token (whether the original token is needed when restoring on a different server, a boolean), and member_name_changed (whether the etcd member name changed between before and after the restore, a boolean).

Notes

What is the datastore of this cluster

In the namespace k3s-ds, create a ConfigMap marker (stage=sqlite). Then write to /root/k3s-ds/datastore.json the fields datastore (sqlite or etcd), db_files (a sorted array of the file names in /var/lib/rancher/k3s/server/db), kine_table (among the table names in state.db, the one kine uses), kine_endpoint (the address after Kine available at in the k3s log), and marker_uid (the UID of marker).

k3s does not attach SQLite directly to the API server; it puts kine, which imitates the etcd API, in between. Check in the log what the API server's --etcd-servers points to. You can see the table names with sqlite3 <파일> .tables (the placeholder is the database file).

You tried the snapshot command but the datastore was SQLite

In the current state, run k3s etcd-snapshot save and save the standard output and standard error together to /root/k3s-ds/snapshot-refused.txt. Then write to /root/k3s-ds/sqlite-backup.txt, one per line, the absolute paths of the directory and file you would have to copy to back up the SQLite datastore.

etcd-snapshot sends a request to the server, and the server takes the snapshot with the embedded etcd. If the datastore is SQLite, the server rejects it and leaves the detailed reason in the server log. SQLite is backed up by copying files without any special command, but the value that encrypts the confidential data in the datastore must be kept along with it.

Move to etcd with a single restart

Write cluster-init: true in /etc/rancher/k3s/config.yaml and restart k3s. Confirm that marker is still alive and write to /root/k3s-ds/migrate.json the fields marker_uid (the UID of marker after the restart), sqlite_file_now (the name the original state.db was changed to in the db directory), node_roles (a sorted array of the node-role.kubernetes.io/ label names of the node, for example ["control-plane"]), and member_name (the content of the db/etcd/name file).

If you start a server that was running on SQLite with cluster-init, k3s moves the SQLite content to etcd. If etcd data is already on disk, this argument is ignored. Look for Migrating content from sqlite to etcd in the log.

Take a named snapshot

Change the stage of marker to snapshotted, and then take a snapshot with k3s etcd-snapshot save --name lab-before. Write to /root/k3s-ds/snapshot.json the fields name (the full name of the snapshot that was created), path (the absolute path of the file), size (in bytes, an integer), and sha256 (the SHA-256 of the file).

--name decides only the first part of the name, and k3s appends the node name and the time. The storage location is the default of --etcd-snapshot-dir. k3s etcd-snapshot ls and kubectl get etcdsnapshotfile show the same snapshots.

You cannot restore with the snapshot alone

Copy the snapshot file from step 4 (with the same name) and the server token file (named token, permission 600) into the /root/k3s-ds/backup/ directory. Write to /root/k3s-ds/token-backup.json the fields token_source (the absolute path of the token original), token_sha256 (the SHA-256 of the token file), and snapshot_sha256 (the SHA-256 of the backed-up snapshot). Do not write the token value itself anywhere.

k3s encrypts the confidential bootstrap data in the datastore with the server token. If you restore with a different token, the snapshot cannot be used. Make sure the permissions do not get wider when you copy.

What gets created after the snapshot

After the snapshot, create a ConfigMap after-snap (x=1) in the namespace k3s-ds and change the stage of marker to changed. Write to /root/k3s-ds/after.json the fields after_uid (the UID of after-snap), after_created (its creationTimestamp), and marker_stage (the current stage of marker).

The point is to see what the restore in the next step does to these two. Write the UID and time exactly so that you can confirm after the restore that this record was genuine.

After the restore, the changes made after the snapshot disappeared

Stop k3s, restore with k3s server --cluster-reset --cluster-reset-restore-path=<4단계 스냅숏 경로> (the placeholder is the path of the snapshot from step 4) (save the entire output to /root/k3s-ds/reset.log), and then start k3s again. Write to /root/k3s-ds/restore.json the fields old_dir (the name of the directory to which the restore moved the previous etcd data), restart_hint (from the sentences in reset.log in which k3s tells you what to do next, the first sentence, up to the part that contains restart without), after_snap_exists (a boolean), marker_stage (the stage of marker after the restore), and member_name (the content of db/etcd/name after the restore).

A restore means running the same binary once separately while the service is stopped. When it finishes it exits by itself and tells you to start again. It does not delete the previous data but moves it aside. A marker file that prevents consecutive resets is created and then deleted after a normal startup.

What remained and what disappeared

Write to /root/k3s-ds/report.json the fields datastore_now (sqlite or etcd), lost_objects (a sorted array of the ConfigMap names that disappeared because of the restore), marker_stage (the current value), snapshot_dir (the absolute path of the directory where the snapshot is stored), restore_needs_token (whether the original token is needed when restoring on a different server, a boolean), and member_name_changed (whether the etcd member name changed between before and after the restore, a boolean).

Write them based on the json files from earlier steps and the current disk and cluster. The grader recalculates the same values from the record files, the etcd-old directory, and the cluster.