Kubernetes Distributions — Build Them Yourself
I tried to SSH in, but there is no shell
Goal
You look into a Talos Linux node that has neither a shell nor SSH using only talosctl and the API, change and revert the machine configuration without a reboot, and distinguish from the logs the mistakes that validation blocks from the ones it cannot block.
Why it matters
k3s, k0s, and kubeadm install Kubernetes on top of Linux, so when a problem occurs, the habit of logging in to the node to read logs and fix files works. Talos has the OS itself rewritten exclusively for Kubernetes, so there is no shell, SSH, or package manager, and the root filesystem is read-only. Instead, each node has a gRPC API (apid), every check and change goes in through that API, and the machine configuration remains as a versioned resource. This design fundamentally blocks the snowflake server of "a setting someone touched only on that node", at the price of making emergency fixes by logging in impossible. So an operator must be able to tell apart, by reading through the API, the case where a configuration is rejected by validation and the case where it passes validation but the service dies. This lab creates both cases on purpose.
Steps
- On the control plane node container
talos-default-controlplane-1, try runningshwithdocker exec, and check whether TCP ports 22 and 50000 of the node IP are open. Write the result to/root/talos-lab/noshell.jsonascontainer,node_ip(the address on the Docker network talos-default),exec_sh_exit(the exit code of docker exec, a number),exec_sh_error(one line of the error message), andport22_openandport50000_open(booleans). - With
talosctl, read the service lists of the two nodes (control plane 10.5.0.2, worker 10.5.0.3) and write to/root/talos-lab/services.jsoncontrolplaneandworker(a sorted array of the service IDs of each node) andcontrolplane_only(a sorted array of the service IDs that exist only on the control plane). - Read the members with
talosctl get membersand write to/root/talos-lab/members.jsontalos_version(the Talos tag of the 10.5.0.2 server, in the form v0.0.0),members(hostname →{"type": 머신 종류, "addresses": 주소 배열}, where the placeholders are the machine type and the array of addresses), andworker_mc_version(the metadata.version of the MachineConfig resourcev1alpha1of worker 10.5.0.3, a number). - Get the kubeconfig from the control plane (10.5.0.2) with
talosctl kubeconfig, save it to/root/talos-lab/kubeconfig(no symbolic links), and using kubectl with that file, write to/root/talos-lab/cluster.jsonapi_server(the server address of that kubeconfig),nodes(node name → InternalIP),kubelet_version,pod_subnets, andservice_subnets(the cluster.network values of the control plane machine configuration), andoverlaps_host(true if either of the two ranges overlaps the host cluster's 10.244.0.0/16 or 10.96.0.0/12). - In
/root/talos-lab/05-labels.yaml, write a strategic merge patch that addslab.talos.dev/pool: blueto the worker'smachine.nodeLabels, and apply it to the worker (10.5.0.3) only, with--mode=no-reboot. Save the output of the apply command (standard output and error) to/root/talos-lab/patch-out.txt, and write to/root/talos-lab/patch.jsonnode,mc_version_beforeandmc_version_after(the worker MachineConfig version before and after applying), andcontainer_started_at(the State.StartedAt of the worker containertalos-default-worker-1). The label must appear on the Kubernetes Node. - Apply the label
lab.talos.dev/canary: "on"to the worker with--mode=try --timeout=30s. Right after applying, read the worker MachineConfig version, confirm that the label appears on the Kubernetes Node, then wait until it reverts and the label disappears, and read the version again. Write to/root/talos-lab/try.jsontimeout_sec(a number),version_during,seen_on_node(a boolean), andversion_after_revert. Thepoollabel from step 5 must remain. - In
/root/talos-lab/07-bad.yaml, write a patch that adds to the workermachine.nodeLabelsa label whose name islab.talos.dev/team name(a name containing a space) and whose value isplatform, and try applying it to the worker. Save the output to/root/talos-lab/rejected.txt, and write the worker MachineConfig version before and after applying to/root/talos-lab/rejected.jsonasversion_beforeandversion_after. - In
/root/talos-lab/08-kubelet.yaml, write a patch that addsmax-pod: "150"to the workermachine.kubelet.extraArgsand apply it to the worker. Right after it is accepted, read the version, look at the kubelet service state and log with talosctl to find the cause, and write to/root/talos-lab/kubelet-diag.jsonaccepted_version(the worker MachineConfig version right after applying),service_state(the STATE of the kubelet service seen then), anderror_line(the one log line stating the cause). Then, in/root/talos-lab/08-fix.yaml, write and apply a patch that removes only that argument with$patch: delete, and make the kubelet healthy and the worker Node Ready again. - Write to
/root/talos-lab/report.jsonshell_in_node(a boolean),ssh_port_open(a boolean),api_port(the Talos API port number),worker_config_changes(the number of times the worker MachineConfig has changed since the first = current version - 1),worker_restarts(the number of times the worker container has restarted since step 5),rejected_field(the configuration path caught by validation in step 7, in the form machine.x),broken_service(the ID of the service that died in step 8),apply_log_line(one line in the worker'stalosctl logs machinedwhere a configuration-apply API call was recorded), andsummary(at least 80 characters on what an immutable OS and API-based management meant in this lab).
Notes
- Two Talos v1.13.10 nodes (control plane 10.5.0.2, worker 10.5.0.3) are running in Docker inside the VM. The talosconfig is already at
/root/.talos/configand the kubeconfig at/root/.kube/config. - Specifying a node:
talosctl -n 10.5.0.3 services— if you omit-n, it fails with something like "nodes are not set". - Reading resources:
talosctl -n <ip> get members,talosctl -n <ip> get mc v1alpha1 -o yaml, and for JSON,-o json | jq -s - Common mistake: putting in a JSON6902 (op/path) patch. This cluster's configuration is multi-document, so it is rejected. Use strategic merge YAML.
- Common mistake: putting in another patch while waiting in try mode. The revert is canceled.
- On this VM,
talosctl dmesgshows the VM's own kernel log (container mode). Look for the cause of a service problem intalosctl logs <서비스>(the placeholder is the service ID). - Bringing up Talos in Docker · Editing machine configuration and apply modes · Configuration patches
You try to log in over SSH, but there is no shell
On the control plane node container talos-default-controlplane-1, try running sh with docker exec, and check whether TCP ports 22 and 50000 of the node IP are open. Write the result to /root/talos-lab/noshell.json as container, node_ip (the address on the Docker network talos-default), exec_sh_exit (the exit code of docker exec, a number), exec_sh_error (one line of the error message), and port22_open and port50000_open (booleans).
The node IP is in NetworkSettings.Networks of docker inspect. To check a port, you can open bash's /dev/tcp/<ip>/<port> together with timeout. The exit code is $? right after the command. Under set -e, capture it in the form || rc=$?.
What is running instead of a shell
With talosctl, read the service lists of the two nodes (control plane 10.5.0.2, worker 10.5.0.3) and write to /root/talos-lab/services.json controlplane and worker (a sorted array of the service IDs of each node) and controlplane_only (a sorted array of the service IDs that exist only on the control plane).
The talosconfig has no default node set, so you must pick a node with -n. Services are also Service resources in the runtime namespace, so you can get them in machine-readable form with talosctl get services -o json. JSON output comes as a series of objects, so gather them with jq -s.
Who are the members of this cluster
Read the members with talosctl get members and write to /root/talos-lab/members.json talos_version (the Talos tag of the 10.5.0.2 server, in the form v0.0.0), members (hostname → {"type": 머신 종류, "addresses": 주소 배열}, where the placeholders are the machine type and the array of addresses), and worker_mc_version (the metadata.version of the MachineConfig resource v1alpha1 of worker 10.5.0.3, a number).
Even if you ask a single node, you get the whole cluster's members (discovery). The server version is the Tag on the Server side of talosctl version. The machine configuration is also a resource, so it can be read with get machineconfig, and the version goes up every time the configuration changes — you will use this value for comparison in later steps.
You get the kubeconfig through the API too
Get the kubeconfig from the control plane (10.5.0.2) with talosctl kubeconfig, save it to /root/talos-lab/kubeconfig (no symbolic links), and using kubectl with that file, write to /root/talos-lab/cluster.json api_server (the server address of that kubeconfig), nodes (node name → InternalIP), kubelet_version, pod_subnets, and service_subnets (the cluster.network values of the control plane machine configuration), and overlaps_host (true if either of the two ranges overlaps the host cluster's 10.244.0.0/16 or 10.96.0.0/12).
You can get the body of the machine configuration as YAML with talosctl get mc v1alpha1 -o jsonpath='{.spec}', and several documents are joined by ---. You can work out whether the ranges overlap with Python's ipaddress overlaps. If you get a kubeconfig twice into the same file, it merges rather than overwrites (-1 is added to the names), so get it only once.
Attach a node label without a reboot
In /root/talos-lab/05-labels.yaml, write a strategic merge patch that adds lab.talos.dev/pool: blue to the worker's machine.nodeLabels, and apply it to the worker (10.5.0.3) only, with --mode=no-reboot. Save the output of the apply command (standard output and error) to /root/talos-lab/patch-out.txt, and write to /root/talos-lab/patch.json node, mc_version_before and mc_version_after (the worker MachineConfig version before and after applying), and container_started_at (the State.StartedAt of the worker container talos-default-worker-1). The label must appear on the Kubernetes Node.
The only way to change the machine configuration is the API. talosctl patch machineconfig (mc for short) fetches the current configuration, merges the patch, and sends it back. Pass the patch file as @파일 (the placeholder is the patch file). This cluster's configuration is multi-document, so the JSON6902 (op/path) format is rejected. You can confirm that there was no reboot by checking whether the container start time is unchanged.
A label that reverted by itself after 30 seconds
Apply the label lab.talos.dev/canary: "on" to the worker with --mode=try --timeout=30s. Right after applying, read the worker MachineConfig version, confirm that the label appears on the Kubernetes Node, then wait until it reverts and the label disappears, and read the version again. Write to /root/talos-lab/try.json timeout_sec (a number), version_during, seen_on_node (a boolean), and version_after_revert. The pool label from step 5 must remain.
Try mode reverts to the previous configuration if there is no other configuration change within a set time after applying. The revert is also a configuration change, so the version goes up once more. If you put in another patch while waiting, the revert is canceled, so be careful.
A patch that validation rejected
In /root/talos-lab/07-bad.yaml, write a patch that adds to the worker machine.nodeLabels a label whose name is lab.talos.dev/team name (a name containing a space) and whose value is platform, and try applying it to the worker. Save the output to /root/talos-lab/rejected.txt, and write the worker MachineConfig version before and after applying to /root/talos-lab/rejected.json as version_before and version_after.
The machine configuration goes through validation before it is written to the node. If it is rejected, nothing should change. --dry-run only shows the merged result and does not validate, so it is not evidence. Offline, you can see the same error by putting the file merged with talosctl machineconfig patch into talosctl validate -m container.
It was accepted, but the kubelet died
In /root/talos-lab/08-kubelet.yaml, write a patch that adds max-pod: "150" to the worker machine.kubelet.extraArgs and apply it to the worker. Right after it is accepted, read the version, look at the kubelet service state and log with talosctl to find the cause, and write to /root/talos-lab/kubelet-diag.json accepted_version (the worker MachineConfig version right after applying), service_state (the STATE of the kubelet service seen then), and error_line (the one log line stating the cause). Then, in /root/talos-lab/08-fix.yaml, write and apply a patch that removes only that argument with $patch: delete, and make the kubelet healthy and the worker Node Ready again.
Validation looks at the format of the configuration document, not whether the kubelet knows that flag. With no shell, look for the cause in talosctl services, in the events of talosctl service <id>, and in talosctl logs <id>. When you remove it, remove just that one key, not all of extraArgs.
How you operated a node without a shell
Write to /root/talos-lab/report.json shell_in_node (a boolean), ssh_port_open (a boolean), api_port (the Talos API port number), worker_config_changes (the number of times the worker MachineConfig has changed since the first = current version - 1), worker_restarts (the number of times the worker container has restarted since step 5), rejected_field (the configuration path caught by validation in step 7, in the form machine.x), broken_service (the ID of the service that died in step 8), apply_log_line (one line in the worker's talosctl logs machined where a configuration-apply API call was recorded), and summary (at least 80 characters on what an immutable OS and API-based management meant in this lab).
Write based on the files you left in earlier steps and the current cluster. Applying a configuration is a gRPC call into machined's MachineService, and every successful call leaves one line in the log. The grader recounts the numbers from the cluster.