TT Lab
Get started
Learn Learning paths Courses

HPC and Slurm

Defining GPUs With GRES

Continue in TT Lab

In one line

In Slurm, a GPU is handled as a generalized device resource called GRES (Generic RESource). Two files must agree with each other — slurm.conf and gres.conf.

Why this was needed

Slurm can count CPUs and memory by itself. But GPUs? FPGAs? NVMe? Slurm does not know these directly, so it created a structure in which the administrator declares them. That is GRES.

How it works

The roles of the two files

File What you write in it Distribution scope
slurm.conf Which GRES types exist, and how many each node has The same on all nodes
gres.conf Which device file that GRES actually corresponds to Can differ per node

On the slurm.conf side:

GresTypes=gpu
NodeName=gpu-node01 ... Gres=gpu:a100:4 State=UNKNOWN

On the gres.conf side:

NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia[0-3] Cores=0-15

The counts of the two must match. If you wrote gpu:a100:8 in slurm.conf but File=/dev/nvidia[0-3] in gres.conf, there are only 4. In this case slurmd leaves gres/gpu count reported lower than configured in the log and DRAINs the node. This is the most common GRES configuration error.

Cores and NUMA affinity

NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia0 Cores=0-15
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia1 Cores=0-15
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia2 Cores=16-31
NodeName=gpu-node01 Name=gpu Type=a100 File=/dev/nvidia3 Cores=16-31

Cores= tells you the CPU cores physically close to that GPU (attached to the same NUMA node). With this information Slurm places the GPU and CPU on the same socket to reduce PCIe round trips. In large-scale training this difference amounts to several percent.

You check the actual topology with nvidia-smi topo -m.

AutoDetect

AutoDetect=nvml

It detects GPUs automatically with the NVML library. It is convenient because it fills in device files and NUMA affinity for you, but NVML must be installed and Slurm must be built with it. And if the auto-detection result differs from the declaration in slurm.conf, the node still drops out — automatic does not take over the verification too.

Manual definition is troublesome, but it is explicit and reproducible. In air-gapped or heterogeneous clusters, manual is often safer.

MIG

If you use MIG on an A100/H100, one physical GPU is split into several instances. In that case each MIG instance becomes a separate GRES entry, and the profile name goes in Type=.

NodeName=gpu-node01 Name=gpu Type=1g.10gb File=/dev/nvidia-caps/...

When the MIG configuration changes, gres.conf must change with it. If you do not automate this, the node drops out every time MIG is reconfigured.

Requesting GPUs

sbatch --gres=gpu:2 train.sh                 # 타입 무관 2장
sbatch --gres=gpu:a100:2 train.sh            # a100 2장
sbatch --gpus=2 train.sh                     # 최신 문법
sbatch --gpus-per-node=2 --nodes=2 train.sh  # 노드당 2장씩 총 4장

Inside a job, CUDA_VISIBLE_DEVICES is set automatically. That value is not the physical device number but a logical number starting from 0. That is why you must not hardcode physical numbers in code.

What it looks like in the field

GPU count mismatch. A GPU disappeared from a node because of a hardware error, but the configuration is unchanged. slurmd reports the count mismatch and the node is DRAINed. This alarm is in fact a good thing — it is better than jobs running on a broken GPU and producing strange results.

Performance does not come out because Cores= was not written. If the GPU and CPU are placed on different sockets, data goes back and forth across PCIe. It stands out especially in multi-GPU training.

What you will do in the next lab

You declare the GRES type and per-node GPUs in slurm.conf, write gres.conf, and go as far as verifying that the counts of the two files match.