FDE Capstone: The Warehouse Got the Same Order Three Times
ImagePullBackOff in an air-gapped cluster
This lab runs on real k3s
A k3s (v1.33.3+k3s1) is running inside the VM, and containerd and the kubelet really pull images. It takes a few minutes to come up the first time. At first the VM can reach the internet (80/443) — step 1 plays the role of the connected import-preparation machine, and in step 2 you yourself block outbound traffic to turn it into the customer air-gapped network. The grading is done by an agent coming in from outside to port 8899 of the VM, so if you block established connections and the loopback, the grading is cut off.
Goal
You fetch images as archives on the connected side and record the hashes, leave ImagePullBackOff as evidence on a k3s whose outbound traffic is blocked, import by two routes, k3s ctr images import and the k3s image directory, confirm even the pull policy that still fails after the import, and leave the import record reconciled against the runtime.
Why it matters
In air-gapped customer sites, half of "the Pod won't come up" is the image. A cluster that does not reach the registry cannot get images by itself, and someone has to put them into the runtime store. But putting them in is not the end. If you put in a file that changed while being carried, nobody knows what is running, and if the manifest is imagePullPolicy: Always or the tag is latest, even with the image you put in, it goes out to ask the registry and fails. An import is not "we copied the file" but proving that "the runtime holds an image whose hash matches, and the Pod uses it".
The expected time is 60 minutes. When the session ends, the VM disappears, so keep any files you want to keep separately before it ends.
Steps
- Fetch the two images
registry.k8s.io/e2e-test-images/busybox:1.36.1-1andregistry.k8s.io/e2e-test-images/nginx:1.15-4as image archives, put them at/root/airgap/bundle/busybox.tarand/root/airgap/bundle/nginx.tar, and writesha256 파일이름 이미지(the placeholders are the file name and the image, and there are two spaces between fields) one per line in/root/airgap/manifest.txt. The image name inside the archive must be exactly the name above. - Write
/root/airgap/egress.sh, which leaves only one set of rules even if you run it again, and block outbound traffic with it. The iptables chain name isAIRGAP-EGRESS, and you send to it once each from OUTPUT and FORWARD. Keep the loopback, established connections, the k3s Pod range (10.42.0.0/16), and the service range (10.43.0.0/16), and block the rest. Write the values measured before and after blocking withcurl -s -o /dev/null -w '%{http_code}' https://registry.k8s.io/v2/into/root/airgap/egress-proof.txtasbefore=andafter=. - In the namespace
airgap, create a Podshop-apithat runssleep 3600with the busybox image (manifest/root/airgap/shop-api.yaml). When the pull fails, saveuid=<파드 UID>(the placeholder is the Pod UID) and the status and events into/root/airgap/pull-failure.txt. - After comparing the hash, import busybox.tar with
k3s ctr images importand save the output into/root/airgap/import-busybox.txt. Makeshop-apiRunning. - Put nginx.tar into the k3s agent image directory (
/var/lib/rancher/k3s/agent/images/) to import it, and make a Podshop-web(manifest/root/airgap/shop-web.yaml) with that image Running. - With the busybox image you imported, create a Pod
shop-always(manifest/root/airgap/shop-always.yaml) withimagePullPolicy: Always, see the failure, and saveuid=andpolicy=and the status into/root/airgap/always.txt. - Keeping the outbound block, recreate
shop-alwayswithimagePullPolicy: IfNotPresentand make it Running. - Leave the import record in
/root/airgap/import-record.json. For each item of theimagesarray, writeref,tar,tar_sha256,imported_via(ctrorimages-dir),runtime_image_id(the id thatk3s crictl inspectireports), andpods(the list of names of the airgap Pods that are Running with that image now), and at the top level writeegress_blockedas true or false.
Notes
- Fetching as an archive:
skopeo copy docker://<이미지> docker-archive:<파일>:<이미지>(the placeholders are the image and the file) - Looking at the runtime store:
k3s ctr -n k8s.io images ls -q, and image details:k3s crictl inspecti <이미지>(the placeholder is the image) - The reason for a Pod's state:
kubectl -n airgap get pod <이름> -o jsonpath='{.status.containerStatuses[0].state.waiting.reason}'(the placeholder is the name) - The kubelet's pull record is also left in the k3s log:
journalctl -u k3s | grep ErrImagePull - Common mistake 1: blocking without keeping established connections (ESTABLISHED) alive. The grading agent's responses are cut off and the grading stops.
- Common mistake 2: not comparing the hash before the import. If you put in a file that changed while being carried, nobody knows what is running.
- Common mistake 3: briefly lifting the block because the import does not seem to work. The customer air-gapped network has no such option.
Build the import bundle on the connected side
Fetch the two images as image archives, put them at /root/airgap/bundle/busybox.tar and /root/airgap/bundle/nginx.tar, and write sha256 파일이름 이미지 (the placeholders are the file name and the image, and there are two spaces between fields) one per line in /root/airgap/manifest.txt.
If you give docker-archive:path:name as the destination of skopeo copy, you get it as a single file with no registry. You have to attach the name so that after the import the Pod finds it under the same name. Measure the hash before carrying it so that the receiving side can compare.
Block outbound traffic to make an air-gapped network
Block outbound traffic with /root/airgap/egress.sh, which leaves only one set of rules even if you run it again, and write the HTTP status codes before and after blocking into /root/airgap/egress-proof.txt as before= and after=.
If you create one new chain and send to it at the very front of OUTPUT and FORWARD, then by emptying and refilling it, it ends up the same no matter how many times you run it. Write what must be kept alive first with RETURN and block at the end. The grading of this VM is a connection coming in from outside to 8899, so if you block established connections, the grading is cut off.
ImagePullBackOff on the blocked cluster
In the namespace airgap, create a Pod shop-api with the busybox image to see the pull failure, and save uid=<파드 UID> (the placeholder is the Pod UID) and the status and events into /root/airgap/pull-failure.txt.
An image with a fixed tag is IfNotPresent if you do not write an imagePullPolicy. It is not in the runtime, so it goes out to pull and gets blocked. The Pod will soon be created again, so leave the UID and the events now — the kubelet's record is also left in the k3s log.
Import directly into the runtime with ctr
Import busybox.tar with k3s ctr images import, save the output into /root/airgap/import-busybox.txt, and make shop-api Running.
The images Kubernetes uses must be in containerd's k8s.io namespace to be seen. Even after the import, the more failures pile up, the longer the kubelet waits before the next attempt (up to 5 minutes), so if you do not want to wait, delete the Pod and create it again with the same manifest.
Import through the k3s image directory
Put nginx.tar into the k3s agent image directory to import it, and make a Pod shop-web with that image Running.
k3s imports a tar placed in the agent's images directory into containerd. Check in the documentation's version table whether it also imports a file put in while running. After waiting a moment until it shows up in the runtime, create the Pod.
A Pod that fails even after the import
With the busybox image you imported, create a Pod shop-always with imagePullPolicy: Always to see the failure, and save uid= and policy= and the status into /root/airgap/always.txt.
Always asks the registry every time, even if the image is local, to resolve the tag to a digest. If it cannot reach the registry, it fails at that step. Also recall that if the tag is latest or absent, the default policy is Always.
Leave the block as it is and fix the policy
Keeping the outbound block, recreate shop-always with imagePullPolicy: IfNotPresent and make it Running.
The container imagePullPolicy of a Pod is a field that cannot be changed after it is created. Lifting the block to get it through is a fix you cannot do in the customer's air-gapped network, so fix the manifest and create it again.
Leave an import record reconciled against the runtime
Leave the records of the two imported images and egress_blocked in /root/airgap/import-record.json.
Do not copy the values of the record by hand; extract them from command output. The runtime image ID is what crictl inspecti tells you, and you gather the Running Pods that use that image from kubectl's JSON. The import route is the method you actually used.