TT Lab
Get started
Learn Learning paths Courses

HPC and Slurm

The Order to Read In When a Node Drops Out

Continue in TT Lab

In one line

Diagnosing a Slurm failure is mostly done by narrowing the scope with sinfo and reading the reason in slurmctld.log.

Why this was needed

When you get a report that "my job isn't running", there are several places to check. Is the job waiting in the queue, has a node dropped out, is authentication broken, is the configuration wrong. But there is an order.

How it works

Step 1 — the overall state

sinfo
sinfo -R                        # DRAIN/DOWN 이유만
sinfo -N -l                     # 노드별 상세

sinfo -R is especially useful. It shows only the nodes that dropped out and their reason strings.

The meaning of node states.

State Meaning
idle Normal, no jobs
alloc / mix All or some resources are allocated
drain An administrator or the system took it out. Does not return automatically
down No response or registration failed
inval The configuration does not match reality
fail / failg A hardware problem was reported

Step 2 — reading the reason

scontrol show node gpu-node01 | grep -i reason
tail -100 /var/log/slurm/slurmctld.log
tail -100 /var/log/slurm/slurmd.log     # 해당 노드에서

Typical messages that appear in the log, and their causes.

Message Cause
Low socket*core*thread count CPUs disagrees with sockets × cores × threads
Low RealMemory The configured memory is larger than the node's actual memory
gres/gpu count reported lower than configured The counts in slurm.conf and gres.conf disagree
Munge decode failed: Expired credential Clock mismatch
Munge decode failed: Invalid credential Key mismatch
partition X has unknown node Y A partition references a node that is not defined
Node unexpectedly rebooted The node registered after a reboot

It is important to tell the two kinds of munge error apart. Expired credential is a clock problem, so you look at NTP, and Invalid credential is a key problem, so you compare /etc/munge/munge.key. They are completely different responses.

Step 3 — when a job is not running

If the nodes are fine but the job only waits, look at the REASON column of squeue.

REASON Response
Resources The requested resources are not yet free. Normal waiting
Priority A higher-priority job is ahead
Dependency Waiting for a predecessor job
PartitionTimeLimit --time exceeds the partition maximum
PartitionNodeLimit The requested number of nodes exceeds the partition limit
ReqNodeNotAvail The specified node is unavailable
QOSMax... A QOS limit applies

Something like PartitionTimeLimit waits forever. Even if resources appear, it never starts, so the user must fix the request and submit again.

Step 4 — recovery

Once you have fixed the cause, bring the node back explicitly.

scontrol update NodeName=gpu-node01 State=RESUME
scontrol update NodeName=gpu-node[01-03] State=RESUME
scontrol reconfigure                     # 설정 변경 반영

If you edited the configuration file, you must run scontrol reconfigure after distributing it to all nodes. If only one node differs, that node drops out again.

Catching nodes dropping out ahead of time

What eats the most time in Slurm operations is not job failures but nodes quietly dropping out. Users only feel that "it got slower", and the cause is that half the cluster is idle.

The reason for dropping out is always recorded.

sinfo -R --format="%50E %12U %19H %N"     # 사유, 지운 사람, 시각, 노드
scontrol show node <노드> | grep -E 'State|Reason|CfgTRES|AllocTRES'

Three common reasons a node becomes drain.

Reason The actual cause
Low RealMemory The memory value in the configuration is larger than the actual. The kernel takes a little
gres/gpu count too low A GPU disappeared (check with nvidia-smi)
Kill task failed The job's processes would not die, so the node was not cleaned up

The first comes up especially often. If you write the physical memory value in slurm.conf as it is, the actual available amount is a bit smaller because of the kernel and reserved areas, and the node drops out right away. Use the value slurmd -C prints and lower it a little further.

slurmd -C          # 이 노드가 보고하는 실제 값

If you changed the configuration, check that it propagated. slurm.conf must be the same on all nodes, and after it changes, scontrol reconfigure is needed. If only one node differs, that node keeps dropping out, and the symptom looks like a problem of that node.

Cleanup failures are usually because processes would not die. If the epilog does not finish, the node is stuck in completing. If it exceeds UnkillableStepTimeout, the node is drained. If this happens often, processes are in the D state waiting on the file system (especially NFS), so the place to fix is not Slurm but the storage.

When you bring it back, clear the reason. If you only restore the state and leave the reason, it drops out again at the next health check.

scontrol update NodeName=<노드> State=RESUME

What it looks like in the field

Clock problems recur periodically. If NTP is dead or a firewall blocks port 123, the clock drifts slowly over several days, and one day munge starts rejecting. It is good to put clock synchronization status into cluster monitoring.

The habit of leaving the DRAIN reason string without clearing it. If you only run scontrol update ... State=RESUME, the reason disappears. Many teams keep the reason somewhere for the investigation record before bringing the node back.

What you will do in the next lab

You receive a fixture containing a broken configuration and logs, identify each of the four causes, create the corrected version, write the recovery procedure, and leave an RCA report.