Writing slurm.conf Precisely
In one line
slurm.conf is a single source of truth that must be identical on every node. If only one node differs, that node quietly drops out.
Why this was needed
A characteristic of Slurm configuration errors is that the error appears late and in an unrelated place. A node drops into INVAL or DOWN, and the log has one line with the cause, but if you do not look at it you wander for days.
How it works
Required keys
ClusterName=labhub-hpc
SlurmctldHost=ctl01
SlurmUser=slurm
SlurmdUser=root
StateSaveLocation=/var/spool/slurmctld
SlurmdSpoolDir=/var/spool/slurmd
SlurmctldPidFile=/run/slurmctld.pid
SlurmdPidFile=/run/slurmd.pid
AuthType=auth/munge
ClusterName— a discriminator when managing several clusters with one DB. It must be lowercase.SlurmctldHost— the controller host name. It must be resolvable on every node by DNS or hosts.SlurmUser— the account that runs slurmctld. Usually a dedicatedslurmaccount.StateSaveLocation— where the queue and job state are stored. If you lose this directory, the information about running jobs disappears. In an HA configuration, you put it on shared storage.
Scheduling and resource selection
SchedulerType=sched/backfill
SelectType=select/cons_tres
SelectTypeParameters=CR_Core_Memory
ProctrackType=proctrack/cgroup
TaskPlugin=task/cgroup,task/affinity
SelectType=select/cons_tres— allocates resources split into cores, memory, and GRES rather than whole nodes. On a GPU cluster it is practically required. The old name iscons_res.SelectTypeParameters=CR_Core_Memory— treats cores and memory both as consumable resources. If you leave memory out, jobs exceeding memory pile onto a single node and OOM occurs.ProctrackType=proctrack/cgroup— tracks a job's processes with cgroup. Without it, child processes remain after the job ends.TaskPlugin=task/cgroup— enforces resource isolation with cgroup. This plugin is what makes GPUs you did not request invisible.
Node definitions
NodeName=gpu-node01 Sockets=2 CoresPerSocket=16 ThreadsPerCore=2 CPUs=64 RealMemory=257000 Gres=gpu:a100:8 State=UNKNOWN
There are two mistakes that occur most often here.
CPUsdiffers fromSockets × CoresPerSocket × ThreadsPerCore. In this case slurmctld puts the node into theINVALstate withLow socket*core*thread count.- You set
RealMemorylarger than the actual physical memory. If it is larger than the value the node reports, the node goes DOWN withLow RealMemory. Conversely, if you set it too small, memory is wasted. The convention is to set it after subtracting the share the OS uses — for a 256GB node, aboutRealMemory=257000(in MB).
You can check a node's actual values with slurmd -C. Pasting that output as it is is the safest.
Partition definitions
PartitionName=batch Nodes=gpu-node[01-03] Default=YES MaxTime=24:00:00 State=UP
PartitionName=short Nodes=gpu-node[01-03] MaxTime=01:00:00 Priority=100 State=UP
- Every node you write in
Nodes=must be defined as a NodeName. If you reference a node that does not exist, apartition ... has unknown node ...warning appears and that node is ignored. - Give
Default=YESto only one partition. - You can use the range notation
gpu-node[01-03].
Logging and state
SlurmctldLogFile=/var/log/slurm/slurmctld.log
SlurmdLogFile=/var/log/slurm/slurmd.log
AccountingStorageType=accounting_storage/none
JobAcctGatherType=jobacct_gather/cgroup
ReturnToService=2
ReturnToService=2 makes a DOWN node return automatically when it registers again with a valid configuration. With the default 0, an administrator must explicitly RESUME it.
After a change
scontrol reconfigure # 대부분의 변경은 이것으로 반영
systemctl restart slurmctld # 일부 핵심 변경은 재기동 필요
scontrol show config | head -40
Do not forget to distribute the configuration file to every node. If only one node differs, that node quietly drops out. That is why most clusters put /etc/slurm on shared storage or distribute it through configuration management.
What it looks like in the field
Problems caused by not using the slurmd -C output as it is. If you calculate by hand and forget ThreadsPerCore or count hyperthreading wrongly, the node drops out as INVAL. Running slurmd -C on that node and pasting the resulting line as it is is the surest way.
What you will do in the next lab
You write a slurm.conf that meets the requirements from scratch, set the munge key permissions, and create a consistency validation script yourself. The grader runs that script against both a valid configuration and a broken one.