Air-Gapped GPU Driver Installation
Building a Repository From the Transfer and Installing
In one line
If you turn the imported .deb bundle into a repository, you can leave installation to apt instead of installing with dpkg one by one in the right order. That is the standard way to import.
Why this was needed
If you install imported files all at once with dpkg -i *.deb, it fails because of the order. You can go one at a time in the right order, but with 30 packages that is work too. And above all, next time you do it again, you have to remember that order.
If you make it a repository, this problem disappears. apt resolves the graph and decides the order. And that repository keeps being reused.
How it works
The procedure
# 1. 반입분을 한곳에 모은다
mkdir -p /srv/airgap/pool && cp /media/usb/*.deb /srv/airgap/pool/
# 2. 색인을 만든다
cd /srv/airgap && dpkg-scanpackages pool /dev/null > pool/Packages
gzip -kf pool/Packages
# 3. 저장소로 등록한다
echo 'deb [trusted=yes] file:/srv/airgap/pool ./' > /etc/apt/sources.list.d/airgap.list
apt-get update
# 4. 먼저 시뮬레이션
apt-get install -s nvidia-driver-550
# 5. 실제 설치
apt-get install -y nvidia-driver-550
Do not skip step 4. -s (simulate) changes nothing and only prints the plan. Here the lines starting with Inst are the packages that will be installed, and if the dependencies do not resolve, The following packages have unmet dependencies appears. This output is the answer to whether the import succeeded.
How to read it when it does not resolve
The following packages have unmet dependencies:
libnvidia-container-tools : Depends: libnvidia-container1 (>= 1.16.2-1) but it is not installable
How to read it: the left side is the one that requires, the right side is what is missing. not installable means that name is not in the repository at all, and but 1.0 is to be installed means it exists but the version does not match. This distinction becomes the basis for making the second import list.
Make the repository information explicit
deb [trusted=yes] file:/srv/airgap/pool ./
[trusted=yes] is a declaration that you trust an unsigned repository. Unless it is a lab or a temporary repository, you must sign. In production, you sign the Release file with your own GPG key and specify the key with [signed-by=/usr/share/keyrings/...].
And if you put a snapshot date in the repository name or path, it is a great help when you investigate six months later.
Checking after installation
dpkg -l 'nvidia*' 'libnvidia*' | grep '^ii' | wc -l
dpkg -l | grep -c '^i[^i]' # ii 가 아닌 것 (문제 있는 상태)
apt-get -f install # 미해결이 있으면 정리
If there is even one state other than ii (iU, iF, and so on), the installation has not completely finished.
The relationship between the driver, CUDA, and the container toolkit
The three are each different things, and their version rules differ too. If you cannot tell them apart, a day goes by on "I installed CUDA but it doesn't work".
[ 호스트 ]
NVIDIA 드라이버 (커널 모듈 + libcuda.so) ← 여기만 커널에 붙는다
nvidia-container-toolkit ← 컨테이너에 장치를 넣어 준다
↓
[ 컨테이너 ]
CUDA 런타임 (libcudart) ← 이미지 안에 들어 있다
cuDNN, PyTorch …
You do not put the driver inside the container. The kernel module is only the host's one, and the toolkit puts /dev/nvidia* and the host's libcuda.so into the container.
The version rule is that the driver must be the same as or higher than CUDA (forward compatibility). With driver 550, up to CUDA 12.4 works. So when you raise the CUDA version of an image, you first check the host driver.
nvidia-smi # 오른쪽 위에 드라이버와 CUDA 상한이 나온다
cat /proc/driver/nvidia/version
If you upgrade the kernel, the module disappears
This is the accident you meet most often. If apt upgrade raises the kernel, there is no NVIDIA module for that kernel and the GPU disappears after a reboot.
# DKMS 가 새 커널에 맞춰 다시 빌드해 준다 — 이것이 설치돼 있어야 한다
dkms status
nvidia/550.90.07, 6.8.0-45-generic, x86_64: installed
# 커널을 고정하는 방법(폐쇄망에서는 이쪽이 안전하다)
apt-mark hold linux-image-generic linux-headers-generic
In an air-gapped network, the headers and compiler needed for the DKMS build must be imported. So pinning the kernel is often the safer choice.
When Secure Boot blocks it
If Secure Boot is on, unsigned kernel modules are not loaded. The symptom is "the installation succeeded but nvidia-smi cannot find the device".
mokutil --sb-state # SecureBoot enabled 이면 이 문제일 수 있다
dmesg | grep -i 'nvidia\|taint\|module verification'
You choose one of three: enroll an MOK key and sign the module, use a driver package signed by the distribution, or turn Secure Boot off. In an air-gapped network, MOK enrollment has to be typed in by a person at the console during the reboot, so if it is an environment with remote access only, you must plan for it in advance.
What it looks like in the field
Design the procedure on the assumption that a second import is inevitable. However well you compute, something drops out of the first import. So mature organizations nail down "simulation after the first import → automatic generation of the shortfall list → second import" as a procedure. If a person makes the shortfall list by hand, something drops out again.
Keep the installation log. Keep the output of apt-get install and /var/log/apt/history.log together with the import record. Later, when you roll back, you can tell which transaction to undo.
What you will do in the next lab
You stand up a repository with the first import batch and make the driver installation succeed, confirm that the toolkit installation fails, and then add the shortfall (the second import batch) and make it succeed again. It is exactly the real air-gapped procedure.