TT Lab
Get started
Learn Learning paths Courses

HPC and Slurm

Writing slurm.conf Precisely

Continue in TT Lab

In one line

slurm.conf is a single source of truth that must be identical on every node. If only one node differs, that node quietly drops out.

Why this was needed

A characteristic of Slurm configuration errors is that the error appears late and in an unrelated place. A node drops into INVAL or DOWN, and the log has one line with the cause, but if you do not look at it you wander for days.

How it works

Required keys

ClusterName=labhub-hpc
SlurmctldHost=ctl01
SlurmUser=slurm
SlurmdUser=root
StateSaveLocation=/var/spool/slurmctld
SlurmdSpoolDir=/var/spool/slurmd
SlurmctldPidFile=/run/slurmctld.pid
SlurmdPidFile=/run/slurmd.pid
AuthType=auth/munge

Scheduling and resource selection

SchedulerType=sched/backfill
SelectType=select/cons_tres
SelectTypeParameters=CR_Core_Memory
ProctrackType=proctrack/cgroup
TaskPlugin=task/cgroup,task/affinity

Node definitions

NodeName=gpu-node01 Sockets=2 CoresPerSocket=16 ThreadsPerCore=2 CPUs=64 RealMemory=257000 Gres=gpu:a100:8 State=UNKNOWN

There are two mistakes that occur most often here.

  1. CPUs differs from Sockets × CoresPerSocket × ThreadsPerCore. In this case slurmctld puts the node into the INVAL state with Low socket*core*thread count.
  2. You set RealMemory larger than the actual physical memory. If it is larger than the value the node reports, the node goes DOWN with Low RealMemory. Conversely, if you set it too small, memory is wasted. The convention is to set it after subtracting the share the OS uses — for a 256GB node, about RealMemory=257000 (in MB).

You can check a node's actual values with slurmd -C. Pasting that output as it is is the safest.

Partition definitions

PartitionName=batch Nodes=gpu-node[01-03] Default=YES MaxTime=24:00:00 State=UP
PartitionName=short Nodes=gpu-node[01-03] MaxTime=01:00:00 Priority=100 State=UP

Logging and state

SlurmctldLogFile=/var/log/slurm/slurmctld.log
SlurmdLogFile=/var/log/slurm/slurmd.log
AccountingStorageType=accounting_storage/none
JobAcctGatherType=jobacct_gather/cgroup
ReturnToService=2

ReturnToService=2 makes a DOWN node return automatically when it registers again with a valid configuration. With the default 0, an administrator must explicitly RESUME it.

After a change

scontrol reconfigure            # 대부분의 변경은 이것으로 반영
systemctl restart slurmctld     # 일부 핵심 변경은 재기동 필요
scontrol show config | head -40

Do not forget to distribute the configuration file to every node. If only one node differs, that node quietly drops out. That is why most clusters put /etc/slurm on shared storage or distribute it through configuration management.

What it looks like in the field

Problems caused by not using the slurmd -C output as it is. If you calculate by hand and forget ThreadsPerCore or count hyperthreading wrongly, the node drops out as INVAL. Running slurmd -C on that node and pasting the resulting line as it is is the surest way.

What you will do in the next lab

You write a slurm.conf that meets the requirements from scratch, set the munge key permissions, and create a consistency validation script yourself. The grader runs that script against both a valid configuration and a broken one.