Defining GPUs With GRES
In one line
In Slurm, a GPU is handled as a generalized device resource called GRES (Generic RESource). Two files must agree with each other — slurm.conf and gres.conf.
Why this was needed
Slurm can count CPUs and memory by itself. But GPUs? FPGAs? NVMe? Slurm does not know these directly, so it created a structure in which the administrator declares them. That is GRES.
How it works
The roles of the two files
| File | What you write in it | Distribution scope |
|---|---|---|
slurm.conf |
Which GRES types exist, and how many each node has | The same on all nodes |
gres.conf |
Which device file that GRES actually corresponds to | Can differ per node |
On the slurm.conf side:
GresTypes=gpu
NodeName=gpu-node01 ... Gres=gpu:a100:4 State=UNKNOWN
On the gres.conf side:
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia[0-3] Cores=0-15
The counts of the two must match. If you wrote gpu:a100:8 in slurm.conf but File=/dev/nvidia[0-3] in gres.conf, there are only 4. In this case slurmd leaves gres/gpu count reported lower than configured in the log and DRAINs the node. This is the most common GRES configuration error.
Cores and NUMA affinity
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia0 Cores=0-15
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia1 Cores=0-15
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia2 Cores=16-31
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia3 Cores=16-31
Cores= tells you the CPU cores physically close to that GPU (attached to the same NUMA node). With this information Slurm places the GPU and CPU on the same socket to reduce PCIe round trips. In large-scale training this difference amounts to several percent.
You check the actual topology with nvidia-smi topo -m.
AutoDetect
AutoDetect=nvml
It detects GPUs automatically with the NVML library. It is convenient because it fills in device files and NUMA affinity for you, but NVML must be installed and Slurm must be built with it. And if the auto-detection result differs from the declaration in slurm.conf, the node still drops out — automatic does not take over the verification too.
Manual definition is troublesome, but it is explicit and reproducible. In air-gapped or heterogeneous clusters, manual is often safer.
MIG
If you use MIG on an A100/H100, one physical GPU is split into several instances. In that case each MIG instance becomes a separate GRES entry, and the profile name goes in Type=.
NodeName=gpu-node01 Name=gpu Type=1g.10gb File=/dev/nvidia-caps/...
When the MIG configuration changes, gres.conf must change with it. If you do not automate this, the node drops out every time MIG is reconfigured.
Requesting GPUs
sbatch --gres=gpu:2 train.sh # 타입 무관 2장
sbatch --gres=gpu:a100:2 train.sh # a100 2장
sbatch --gpus=2 train.sh # 최신 문법
sbatch --gpus-per-node=2 --nodes=2 train.sh # 노드당 2장씩 총 4장
Inside a job, CUDA_VISIBLE_DEVICES is set automatically. That value is not the physical device number but a logical number starting from 0. That is why you must not hardcode physical numbers in code.
What it looks like in the field
GPU count mismatch. A GPU disappeared from a node because of a hardware error, but the configuration is unchanged. slurmd reports the count mismatch and the node is DRAINed. This alarm is in fact a good thing — it is better than jobs running on a broken GPU and producing strange results.
Performance does not come out because Cores= was not written. If the GPU and CPU are placed on different sockets, data goes back and forth across PCIe. It stands out especially in multi-GPU training.
What you will do in the next lab
You declare the GRES type and per-node GPUs in slurm.conf, write gres.conf, and go as far as verifying that the counts of the two files match.