TT Lab
Get started
Learn Learning paths Courses

Building clusters with Kubespray and Terraform

Air-gapped installs are a list problem — pin down what Kubespray fetches while you're still outside

Continue in TT Lab

One-line summary

An air-gapped installation of kubespray is the job of "making a list of what to download → setting it up internally under the same paths → changing only the addresses with inventory variables", and this module covers up to the list and the variables.

Why this was needed

kubespray downloads quite a lot from the internet during installation. In this course's measured installation too, it downloaded github.com releases (containerd, runc, etcd, the CNI plugins, calicoctl), dl.k8s.io (kubeadm, kubelet, kubectl), images from registry.k8s.io, quay.io, and docker.io, and packages from the Ubuntu apt repository. The servers at financial, public-sector, and manufacturing sites cannot reach any of these. So the air-gapped section of the kubespray documentation divides what you must bring in advance into five — static files (binaries and archives), OS packages, container images, and optionally Python packages and Helm charts. And it lists what you must set up internally to match — an HTTP mirror for files, an internal deb and rpm repository, an internal container registry, and optionally PyPI and a Helm repository.

How it works

The tool makes the list. contrib/offline/generate_list.sh extracts *_download_url and the image repos and tags from download.yml to make a template, and fills in the variables with a small playbook (generate_list.yml) to produce temp/files.list and temp/images.list. Measured on this VM, it produces 24 file lines and 48 image lines in 3.6 seconds. It contains even those of CNIs that are not turned on (cilium, flannel, and so on) and of add-ons, so it is more generous than what an actual installation uses — if you want to finish the import review in one pass, the generous side is better. The image sources were 18 from quay.io, 18 from registry.k8s.io, 10 from docker.io, and 2 from ghcr.io.

A trap: the list playbook runs on localhost. The README tells you to write the version variables in the inventory or group_vars and pass them with -i. But the target of generate_list.yml is hosts: localhost, and localhost does not belong to any group in the inventory, so it receives only the group_vars of all. If, as in this course, you write kube_version in group_vars/k8s_cluster (where the kubespray sample puts it), the list playbook does not see that value and makes the list with the defaults. In measurement the inventory was 1.35.8 but the list came out as 1.36.4, with neither an error nor a warning. You must give -e kube_version=1.35.8 or put the version in group_vars/all.

You change the addresses with variables. The variables shown by the sample's group_vars/all/offline.yml and the documentation fall in two branches.

이미지   kube_image_repo · gcr_image_repo · docker_image_repo · quay_image_repo · github_image_repo  → "{{ registry_host }}"
파일     github_url · dl_k8s_io_url · storage_googleapis_url · get_helm_url                         → "{{ files_repo }}/<원래 도메인>"
노드     containerd_registries_mirrors(containerd 2) · containerd_registry_auth(1.7)                 → 사내 레지스트리를 믿게
OS 패키지 ubuntu_repo · debian_repo · yum_repo                                                         → 사내 저장소

For images, only the registry address changes and the path (coredns/coredns:v…) stays the same. For files too, if you put the original domain as the first directory as the documentation's tip says (files_repo/dl.k8s.io/release/…), you can move mechanically from the original URL to the mirror URL. The reason to put these variables in all is that nodes that handle only etcd also have to download, and the list playbook above also has to read them. In measurement, with the offline variables placed in all, all 72 lines of the list were changed to internal addresses without a single one missed.

The control node is also something to bring in. kubespray v2.32.0 requires ansible==12.3.0 and a few Python packages in exact versions. To run pip install -r requirements.txt on a control node in an air-gapped network, you have to download the wheels beforehand and bring them, and the Python version and architecture of the environment where you downloaded and the one where you install must be the same. You can first check outside whether it resolves without the internet with pip install --dry-run --no-index --find-links.

What it looks like in the field

Air-gapped import usually goes "submit list → review → import → install", and if the list is wrong, the whole schedule slips by one round. The three most common mistakes are these. A list for a different version (the trap above), offline variables that missed one image address (only that one image tries to download from the original domain and stops), and forgetting the control node's Ansible. The lab in this module is arranged in the order of checking these three in advance while still outside.

Actually starting an internal registry and moving the images on the list into it (manage-offline-container-images.sh), serving the file mirror with nginx (manage-offline-files.sh), and running the installation to the end on a node cut off from the internet continue in a course that covers air-gapped networks separately. The list, the mapping table, and the Python bundle you make here are the input of that course.

What you will do in the next lab

You make the default list with generate_list.sh and count by source. You check how the versions in the list differ when you pass the inventory versus when you pass the version with -e. You fill offline.yml with the internal mirror addresses to see whether the list points entirely to the internal network, and make a mapping table between the originals and the mirror. You add the containerd configuration that makes nodes use the internal registry, and finally check that the Python bundle, which includes Ansible, resolves without the internet.