Defining GPUs as GRES
Goal
You write slurm.conf and gres.conf together to define GPUs as GRES, and verify that the counts of the two files match.
Why it matters
The number one GRES configuration error is a mismatch between the GPU count in slurm.conf and the device count in gres.conf. If you wrote gpu:a100:8 in slurm.conf but File=/dev/nvidia[0-3] in gres.conf, there are only 4, and slurmd leaves count reported lower than configured and DRAINs the node. This check is a typical item that must be automated with a script before deployment.
And if you tell Slurm the NUMA affinity with Cores=, it places the GPU and CPU on the same socket to reduce PCIe round trips. In large-scale training this difference amounts to several percent.
You continue using the /etc/slurm/slurm.conf you made in the previous lab (slurm-conf).
Steps
- Add
GresTypes=gputo/etc/slurm/slurm.conf. - Add
Gres=gpu:a100:4to theNodeNamelines of the three nodes. - Create
/etc/slurm/gres.conf. For each of the three nodes there must be a line that includesNodeName=,Name=gpu,Type=a100, andFile=, andFilemust point to 4 devices. - For
gpu-node01, write the devices one by one on four lines, and specifyCores=0-15for/dev/nvidia0and/dev/nvidia1, andCores=16-31for/dev/nvidia2and/dev/nvidia3. (For the other two nodes you may leave the range notation of step 3 as it is.) - Add a GPU-only partition
gputo/etc/slurm/slurm.conf.Nodes=is all three nodes,MaxTime=72:00:00, andState=UP. - Make
/root/slurm/gres-check.txtwith the following two lines. The values are the actual counts obtained by parsing the configuration files.SLURM_CONF_GPUS=<slurm.conf 의 Gres 로 선언된 GPU 총 개수>/GRES_CONF_DEVICES=<gres.conf 의 File 이 가리키는 장치 총 개수>(the placeholders are the total number of GPUs declared as Gres in slurm.conf and the total number of devices the File entries in gres.conf point to) The two values must be equal. - Make
/root/slurm/autodetect.txtwith the following 3 lines.AUTODETECT=nvml/NEEDS=nvml-library/MANUAL_PRO=reproducible - Make
/root/slurm/gres-report.txtwith the following 4 lines.GRESTYPES=gpu/NODES=3/GPUS_PER_NODE=4/PARTITIONS=<slurm.conf 의 PartitionName 줄 수>(the placeholder is the number of PartitionName lines in slurm.conf)
Notes
File=/dev/nvidia[0-3]means 4 devices. You must expand the range notation to count.- You check the actual topology with
nvidia-smi topo -m. - Get the count in step 6 by parsing with the shell or python3. You can count by hand too, but if you do it with a script, it is reused up to step 8.
- Common mistake 1: fixing only one node in step 2 and leaving out the rest. All three nodes are needed.
- Common mistake 2: in step 4, splitting the devices into four lines but not deleting the range-notation line from step 3, so that the devices of
gpu-node01are counted as 8.
Declare the GRES type
Add GresTypes=gpu to /etc/slurm/slurm.conf.
One line in the top area of slurm.conf is enough. Separate multiple types with commas.
Declare the GPUs per node
Add Gres=gpu:a100:4 to the NodeName lines of the three nodes.
Add Gres= to the NodeName line. The format is name:type:count.
Write gres.conf
Create /etc/slurm/gres.conf. For each of the three nodes there must be a line that includes NodeName=, Name=gpu, Type=a100, and File=, and File must point to 4 devices.
You can write one line per node or one line per device. Using the range notation is concise.
Write the NUMA affinity
For gpu-node01, write the devices one by one on four lines, and specify Cores=0-15 for /dev/nvidia0 and /dev/nvidia1, and Cores=16-31 for /dev/nvidia2 and /dev/nvidia3. (For the other two nodes you may leave the range notation of step 3 as it is.)
Write the core range close to each GPU. Split the devices in half and distribute them across the two sockets.
GPU-only partition
Add a GPU-only partition gpu to /etc/slurm/slurm.conf. Nodes= is all three nodes, MaxTime=72:00:00, and State=UP.
Define one more, separate from the existing partitions. Set the maximum time long.
Verify that the counts match
Make /root/slurm/gres-check.txt with the following two lines. The values are the actual counts obtained by parsing the configuration files.
SLURM_CONF_GPUS=<slurm.conf 의 Gres 로 선언된 GPU 총 개수> / GRES_CONF_DEVICES=<gres.conf 의 File 이 가리키는 장치 총 개수> (the placeholders are the total number of GPUs declared as Gres in slurm.conf and the total number of devices the File entries in gres.conf point to)
The two values must be equal.
Compare the sum of the Gres counts in slurm.conf with the number of devices that come out of the File ranges in gres.conf.
AutoDetect comparison note
Make /root/slurm/autodetect.txt with the following 3 lines.
AUTODETECT=nvml / NEEDS=nvml-library / MANUAL_PRO=reproducible
Summarize the tradeoff between automatic detection and manual definition as key=value.
GRES design report
Make /root/slurm/gres-report.txt with the following 4 lines.
GRESTYPES=gpu / NODES=3 / GPUS_PER_NODE=4 / PARTITIONS=<slurm.conf 의 PartitionName 줄 수> (the placeholder is the number of PartitionName lines in slurm.conf)
Get the values by parsing the configuration you wrote.