Writing slurm.conf
Goal
You write a slurm.conf that meets the requirements from scratch, and create a consistency validation script yourself.
Why it matters
Slurm configuration errors appear late and in unrelated places. A node drops into INVAL or DOWN, and the log has one line with the cause, but if you do not look at it you wander for days. The two most common are CPUs differing from sockets × cores × threads and a partition referencing a node that is not defined. Both are mistakes that occur when a person calculates by hand, and both can be caught in advance by a script.
The actual slurmd is not run. The configuration you write and the validation script are graded.
Requirements of this lab
- Cluster name
labhub-hpc, controller hostctl01 - 3 nodes:
gpu-node01togpu-node03, each with 2 sockets / 16 cores per socket / 2 threads per core / 257000MB of memory / 4 A100 GPUs - 2 partitions:
batch(all nodes, default, maximum 24 hours) andshort(all nodes, maximum 1 hour)
Steps
- Create the two directories
/etc/slurmand/root/slurm. - Create
/etc/slurm/slurm.confand put in the following keys.ClusterName=labhub-hpc,SlurmctldHost=ctl01,SlurmUser=slurm,StateSaveLocation=/var/spool/slurmctld,SlurmdSpoolDir=/var/spool/slurmd,AuthType=auth/munge - Add the scheduling-related keys.
SchedulerType=sched/backfill,SelectType=select/cons_tres,SelectTypeParameters=CR_Core_Memory,ProctrackType=proctrack/cgroup,TaskPlugin=task/cgroup,task/affinity - Define the 3 nodes. Each line must start with
NodeName=and includeSockets,CoresPerSocket,ThreadsPerCore,CPUs,RealMemory, andState. TheCPUsvalue must match sockets × cores × threads. - Define the 2 partitions.
batchhasDefault=YESandMaxTime=24:00:00, andshorthasMaxTime=01:00:00. In both partitions, the nodes written inNodes=must be the ones defined in step 4. - Add the log and accounting keys.
SlurmctldLogFile=/var/log/slurm/slurmctld.log,SlurmdLogFile=/var/log/slurm/slurmd.log,JobAcctGatherType=jobacct_gather/cgroup,ReturnToService=2 - Create the
mungeuser and group, generate/etc/munge/munge.keyas 1024 bytes of random data, and set the permission0400and the ownermunge:munge. - Write
/root/slurm/validate.sh. It takes the slurm.conf path as its first argument, checks the following two rules, and must exit with code 0 if there is no violation and with a nonzero value if there is.- Rule 1: On every
NodeName=line,CPUsmust equalSockets × CoresPerSocket × ThreadsPerCore - Rule 2: The nodes that appear in
Nodes=on everyPartitionName=line must be defined withNodeName=(expand and check the range notationgpu-node[01-03]) The grader runs it against both your/etc/slurm/slurm.conf(which must pass) and/opt/fixtures/slurm/broken/slurm.conf(which must fail).
- Rule 1: On every
Notes
- Generating a random key:
dd if=/dev/urandom of=/etc/munge/munge.key bs=1 count=1024 - Expanding the range notation:
gpu-node[01-03]→gpu-node01 gpu-node02 gpu-node03. You can handle it with bash brace expansion or seq. - The actual values of a node come out as they are if you run
slurmd -Con that node. It is safer than calculating by hand. - Common mistake 1: getting the calculation wrong in step 4, such as
CPUs=32. 2×16×2 = 64. - Common mistake 2: a step 8 script that does not expand the range notation and so judges even a valid configuration as a failure.
Prepare the configuration directories
Create the two directories /etc/slurm and /root/slurm.
The standard path is /etc/slurm. Also create a separate working directory.
Write the basic keys
Create /etc/slurm/slurm.conf and put in the following keys.
ClusterName=labhub-hpc, SlurmctldHost=ctl01, SlurmUser=slurm, StateSaveLocation=/var/spool/slurmctld, SlurmdSpoolDir=/var/spool/slurmd, AuthType=auth/munge
The basics are the cluster name, the controller host, the run account, the state save path, and the authentication method.
Scheduler and resource selection
Add the scheduling-related keys.
SchedulerType=sched/backfill, SelectType=select/cons_tres, SelectTypeParameters=CR_Core_Memory, ProctrackType=proctrack/cgroup, TaskPlugin=task/cgroup,task/affinity
On a GPU cluster you need a selection plugin that splits resources per core. Memory must also be included among the consumable resources.
Define the nodes
Define the 3 nodes. Each line must start with NodeName= and include Sockets, CoresPerSocket, ThreadsPerCore, CPUs, RealMemory, and State. The CPUs value must match sockets × cores × threads.
CPUs must match sockets × cores × threads. Check the calculation twice.
Define the partitions
Define the 2 partitions. batch has Default=YES and MaxTime=24:00:00, and short has MaxTime=01:00:00. In both partitions, the nodes written in Nodes= must be the ones defined in step 4.
Every node written in Nodes must be defined. Specify only one default partition.
Logs and accounting
Add the log and accounting keys.
SlurmctldLogFile=/var/log/slurm/slurmctld.log, SlurmdLogFile=/var/log/slurm/slurmd.log, JobAcctGatherType=jobacct_gather/cgroup, ReturnToService=2
You need log paths for the controller and for the nodes. Also specify the resource-gathering plugin.
Prepare the munge key
Create the munge user and group, generate /etc/munge/munge.key as 1024 bytes of random data, and set the permission 0400 and the owner munge:munge.
munged comes up only if the key file's permission and owner are exact. There is also a prescribed size.
Consistency validation script
Write /root/slurm/validate.sh. It takes the slurm.conf path as its first argument, checks the following two rules, and must exit with code 0 if there is no violation and with a nonzero value if there is.
- Rule 1: On every
NodeName=line,CPUsmust equalSockets × CoresPerSocket × ThreadsPerCore - Rule 2: The nodes that appear in
Nodes=on everyPartitionName=line must be defined withNodeName=(expand and check the range notationgpu-node[01-03]) The grader runs it against both your/etc/slurm/slurm.conf(which must pass) and/opt/fixtures/slurm/broken/slurm.conf(which must fail).
The grader runs it against both the valid configuration and the broken configuration in the fixture. Distinguish them by the exit code.