TT Lab
Get started
Learn Learning paths Courses

Air-Gapped GPU Driver Installation

The Verification Plan and the Rollback Path

Continue in TT Lab

In one line

In air-gapped work, you do not make changes that have no rollback path. That is because going back out to download something is impossible.

Why this was needed

On a connected network, if an installation goes wrong you can just download again. An air-gapped network is different. If you deleted something and want it back, you have to go through import review again. One irreversible operation costs several days.

So half of an air-gapped work procedure is "recording what you changed and deciding in advance how to undo it".

How it works

The verification plan

In an environment with an actual GPU, installation verification goes in this order.

# 1. 커널 층
lsmod | grep -c nvidia
ls /dev/nvidia*

# 2. 유저 공간 층
nvidia-smi
nvidia-smi --query-gpu=name,driver_version,memory.total --format=csv

# 3. 컨테이너 주입 층
nvidia-ctk cdi list
nvidia-container-cli info

# 4. 실제 컨테이너
podman run --rm --device nvidia.com/gpu=all <cuda 이미지> nvidia-smi

Checking the layers in order is the key. If you try only step 4 and it fails, you do not know which layer is the problem. If you go up from step 1, the point of failure is the cause.

Package integrity

Even after installation, you can check that the files are unchanged.

dpkg -V nvidia-driver-550        # 설치 시점의 해시와 비교
dpkg -V | head                   # 전체 검사

The output format is a string such as ??5??????. 5 is an md5 mismatch (the content changed), and c marks a configuration file. A changed configuration file is normal, but if a binary changed, you must investigate.

Version pinning

For the driver in particular, pinning matters. When the kernel goes up, a DKMS rebuild is needed, and in an air-gapped network that is yet another import.

apt-mark hold nvidia-driver-550 nvidia-container-toolkit
apt-mark hold linux-image-generic linux-headers-generic
apt-mark showhold

Pinning even the kernel together is the practice for GPU nodes.

The rollback path

apt does not directly support transaction rollback (there is nothing like dnf's history undo). So you prepare it manually.

# 설치 전 상태 기록
dpkg -l > /var/log/labhub/before.txt
apt-mark showhold > /var/log/labhub/holds.txt

# 설치 로그 확인
cat /var/log/apt/history.log
grep -A3 'Start-Date' /var/log/apt/history.log | tail -20

# 롤백 (제거)
apt-get purge -y nvidia-container-toolkit
apt-get autoremove -y

/var/log/apt/history.log records each transaction's start time, command line, and installed and removed packages. This file is in effect the transaction log.

A surer rollback is a system snapshot. If you take an LVM snapshot or a VM snapshot before installation, you can go back from any failure. Many organizations do this before work on air-gapped GPU nodes.

What must not be rolled back

Downgrading base packages such as the kernel or glibc is dangerous. They sit at the root of the dependency graph, so the moment you roll them back, the whole system can become unstable. Such things are not rollback targets but targets that should be raised carefully in the first place.

What it looks like in the field

You removed the driver and X does not come up. If you purge nvidia-driver-*, the display driver disappears with it. It does not matter on a headless server, but on a workstation that needs a console, you must check the nouveau fallback.

You put a hold on and forget it. Months later, the reason security patches are not applied is that hold. You need a procedure for reviewing the hold list regularly, and you must record why you placed it.

Splitting verification into layers

The sentence "the GPU works" includes at least four layers. You can tell where it broke only by checking one by one from the bottom up.

4. 프레임워크    torch.cuda.is_available() == True, 실제 연산이 맞는 값을 낸다
3. 컨테이너      파드 안에서 nvidia-smi 가 GPU 를 본다
2. 스케줄러      노드에 nvidia.com/gpu 자원이 광고되어 있다
1. 호스트        nvidia-smi 가 장치를 나열하고 모듈이 로드돼 있다

The check commands for each layer.

# 1. 호스트
nvidia-smi && lsmod | grep nvidia

# 2. 스케줄러 — 자원이 0 이면 device plugin 이 안 도는 것
kubectl get nodes -o custom-columns=N:.metadata.name,GPU:.status.allocatable.'nvidia\.com/gpu'

# 3. 컨테이너
kubectl run gpu-check --rm -it --restart=Never \
  --image=nvidia/cuda:12.4.0-base-ubuntu22.04 \
  --limits=nvidia.com/gpu=1 -- nvidia-smi

# 4. 프레임워크 — 그리고 값이 맞는지까지
python -c "import torch; a=torch.randn(1000,1000,device='cuda'); print((a@a).sum().item())"

In step 4, checking even the value is important. There are cases where is_available() is True but the actual computation gives NaN (a driver and CUDA version mismatch, an ECC error).

Preparing the rollback path in advance

In an air-gapped network, "download it again from the internet" is impossible, so you start by securing first what you will go back to.

# 되돌리기
apt-get install --allow-downgrades nvidia-driver-550=550.90.07-0ubuntu1
update-initramfs -u && reboot

If you leave out update-initramfs, the old module is still loaded after the reboot and you get "I rolled back but nothing changed".

A format for recording verification results

## GPU 설치 검증 2026-09-06 (node-gpu-03)

| 층 | 확인 | 결과 |
|---|---|---|
| 호스트 | nvidia-smi | 550.90.07, A100 ×2 인식 |
| 스케줄러 | allocatable | nvidia.com/gpu: 2 |
| 컨테이너 | 파드 안 nvidia-smi | 정상 |
| 프레임워크 | torch matmul | 값 일치, 3.2 TFLOPS |

- 되돌리기 검증: 550.54 로 다운그레이드 후 재부팅 → 정상, 다시 550.90.07 로 복귀
- 남은 위험: 커널 자동 업데이트 (apt-mark hold 로 고정함)

What you will do in the next lab

You create an integrity manifest of the imported items and verify it, and reproduce tampering to confirm that verification fails. You pin versions, check the installation transaction record, and actually perform a rollback.