The Order to Read In When a Node Drops Out
In one line
Diagnosing a Slurm failure is mostly done by narrowing the scope with sinfo and reading the reason in slurmctld.log.
Why this was needed
When you get a report that "my job isn't running", there are several places to check. Is the job waiting in the queue, has a node dropped out, is authentication broken, is the configuration wrong. But there is an order.
How it works
Step 1 — the overall state
sinfo
sinfo -R # DRAIN/DOWN 이유만
sinfo -N -l # 노드별 상세
sinfo -R is especially useful. It shows only the nodes that dropped out and their reason strings.
The meaning of node states.
| State | Meaning |
|---|---|
idle |
Normal, no jobs |
alloc / mix |
All or some resources are allocated |
drain |
An administrator or the system took it out. Does not return automatically |
down |
No response or registration failed |
inval |
The configuration does not match reality |
fail / failg |
A hardware problem was reported |
Step 2 — reading the reason
scontrol show node gpu-node01 | grep -i reason
tail -100 /var/log/slurm/slurmctld.log
tail -100 /var/log/slurm/slurmd.log # 해당 노드에서
Typical messages that appear in the log, and their causes.
| Message | Cause |
|---|---|
Low socket*core*thread count |
CPUs disagrees with sockets × cores × threads |
Low RealMemory |
The configured memory is larger than the node's actual memory |
gres/gpu count reported lower than configured |
The counts in slurm.conf and gres.conf disagree |
Munge decode failed: Expired credential |
Clock mismatch |
Munge decode failed: Invalid credential |
Key mismatch |
partition X has unknown node Y |
A partition references a node that is not defined |
Node unexpectedly rebooted |
The node registered after a reboot |
It is important to tell the two kinds of munge error apart. Expired credential is a clock problem, so you look at NTP, and Invalid credential is a key problem, so you compare /etc/munge/munge.key. They are completely different responses.
Step 3 — when a job is not running
If the nodes are fine but the job only waits, look at the REASON column of squeue.
| REASON | Response |
|---|---|
Resources |
The requested resources are not yet free. Normal waiting |
Priority |
A higher-priority job is ahead |
Dependency |
Waiting for a predecessor job |
PartitionTimeLimit |
--time exceeds the partition maximum |
PartitionNodeLimit |
The requested number of nodes exceeds the partition limit |
ReqNodeNotAvail |
The specified node is unavailable |
QOSMax... |
A QOS limit applies |
Something like PartitionTimeLimit waits forever. Even if resources appear, it never starts, so the user must fix the request and submit again.
Step 4 — recovery
Once you have fixed the cause, bring the node back explicitly.
scontrol update NodeName=gpu-node01 State=RESUME
scontrol update NodeName=gpu-node[01-03] State=RESUME
scontrol reconfigure # 설정 변경 반영
If you edited the configuration file, you must run scontrol reconfigure after distributing it to all nodes. If only one node differs, that node drops out again.
Catching nodes dropping out ahead of time
What eats the most time in Slurm operations is not job failures but nodes quietly dropping out. Users only feel that "it got slower", and the cause is that half the cluster is idle.
The reason for dropping out is always recorded.
sinfo -R --format="%50E %12U %19H %N" # 사유, 지운 사람, 시각, 노드
scontrol show node <노드> | grep -E 'State|Reason|CfgTRES|AllocTRES'
Three common reasons a node becomes drain.
| Reason | The actual cause |
|---|---|
Low RealMemory |
The memory value in the configuration is larger than the actual. The kernel takes a little |
gres/gpu count too low |
A GPU disappeared (check with nvidia-smi) |
Kill task failed |
The job's processes would not die, so the node was not cleaned up |
The first comes up especially often. If you write the physical memory value in slurm.conf as it is, the actual available amount is a bit smaller because of the kernel and reserved areas, and the node drops out right away. Use the value slurmd -C prints and lower it a little further.
slurmd -C # 이 노드가 보고하는 실제 값
If you changed the configuration, check that it propagated. slurm.conf must be the same on all nodes, and after it changes, scontrol reconfigure is needed. If only one node differs, that node keeps dropping out, and the symptom looks like a problem of that node.
Cleanup failures are usually because processes would not die. If the epilog does not finish, the node is stuck in completing. If it exceeds UnkillableStepTimeout, the node is drained. If this happens often, processes are in the D state waiting on the file system (especially NFS), so the place to fix is not Slurm but the storage.
When you bring it back, clear the reason. If you only restore the state and leave the reason, it drops out again at the next health check.
scontrol update NodeName=<노드> State=RESUME
What it looks like in the field
Clock problems recur periodically. If NTP is dead or a firewall blocks port 123, the clock drifts slowly over several days, and one day munge starts rejecting. It is good to put clock synchronization status into cluster monitoring.
The habit of leaving the DRAIN reason string without clearing it. If you only run scontrol update ... State=RESUME, the reason disappears. Many teams keep the reason somewhere for the investigation record before bringing the node back.
What you will do in the next lab
You receive a fixture containing a broken configuration and logs, identify each of the four causes, create the corrected version, write the recovery procedure, and leave an RCA report.