Diagnosing a Broken Cluster
Goal
You receive a broken Slurm configuration and logs, identify each of the four causes, fix them, and leave a recovery procedure and an RCA report.
Why it matters
A characteristic of Slurm failures is that a single log line tells you the whole cause. The problem is that if you do not know the order for finding that one line, it takes days. The order is fixed — narrow the scope with sinfo, read the reason in slurmctld.log, and confirm that line in the configuration file.
It is especially important to tell the two kinds of munge error apart. Expired credential is a clock problem, so you look at NTP, and Invalid credential is a key problem, so you compare the key files. They are completely different responses.
The fixtures are in /opt/fixtures/slurm/broken/. They are four files: slurm.conf, gres.conf, slurmctld.log, and sinfo.txt.
Steps
- Create the
/root/rcadirectory, and write, one per line, the distinct node names in a non-normal state (not idle/alloc/mix) fromsinfo.txtin/root/rca/bad-nodes.txt. - Find the node definition mismatch and write it in
/root/rca/node.txtas two lines.NODE=<문제 노드 이름>/KEY=<불일치가 난 설정 키 이름>(the placeholders are the problem node name and the name of the configuration key that mismatched) - Identify the munge authentication failure and write it in
/root/rca/munge.txtas two lines.KIND=<expired|invalid>/CAUSE=<clock|key> - Write, in one line in
/root/rca/partition.txt, the node name that a partition references but that is not defined. - Write the GRES count mismatch in
/root/rca/gres.txtas two lines.CONFIGURED=<slurm.conf 가 선언한 그 노드의 GPU 수>/REPORTED=<gres.conf 가 실제로 가리키는 장치 수>(the placeholders are the number of GPUs slurm.conf declares for that node and the number of devices gres.conf actually points to) - Submit
/root/rca/slurm.confand/root/rca/gres.confwith all four problems fixed. Make the problem node'sCPUsmatch reality, remove the nonexistent node from the partition's node list, and make the GRES counts match. (munge is an operations problem, not a configuration file problem, so it is not dealt with in these files.) - Write
/root/rca/recovery.sh. It must contain the following two kinds of commands, in order.- The command that applies configuration changes (
scontrol reconfigure) - The command that brings the dropped nodes back (
scontrol update NodeName=... State=RESUME)
- The command that applies configuration changes (
- Make
/root/rca/report.txtwith the following 6 lines.CAUSE1=cpus-mismatch/CAUSE2=munge-<3번의 CAUSE 값>/CAUSE3=unknown-node/CAUSE4=gres-count/BAD_NODES=<1번 파일의 줄 수>/FIXED=yes(the placeholders are the CAUSE value from step 3 and the number of lines in the file from step 1)
Notes
- To see only errors in the log,
grep -i error /opt/fixtures/slurm/broken/slurmctld.logis convenient. - The range notation
gpu-node[01-04]means 4 nodes. Expand it to compare. - It is good to check the step 6 files with the
validate.shfrom the previous lab. - Common mistake 1: in step 1, counting a node twice because the same node appears in several partitions. Write only distinct node names.
- Common mistake 2: in step 6, raising the GRES count on the gres.conf side. The actual devices number 4, so matching the slurm.conf side to 4 is right.
Summarize the symptoms
Create the /root/rca directory, and write, one per line, the distinct node names in a non-normal state (not idle/alloc/mix) from sinfo.txt in /root/rca/bad-nodes.txt.
Count the nodes in a non-normal state in the sinfo output. Write the state strings exactly.
Find the node definition mismatch
Find the node definition mismatch and write it in /root/rca/node.txt as two lines.
NODE=<문제 노드 이름> / KEY=<불일치가 난 설정 키 이름> (the placeholders are the problem node name and the name of the configuration key that mismatched)
The log says which node dropped out and why. Confirm that node's line in the configuration file.
Identify the authentication failure
Identify the munge authentication failure and write it in /root/rca/munge.txt as two lines.
KIND=<expired|invalid> / CAUSE=<clock|key>
There are two kinds of munge error. Look at the wording in the log as it is and decide which one it is.
Reference to an undefined node
Write, in one line in /root/rca/partition.txt, the node name that a partition references but that is not defined.
Compare the partition's node list with the nodes actually defined. You must expand the range notation.
GRES count mismatch
Write the GRES count mismatch in /root/rca/gres.txt as two lines.
CONFIGURED=<slurm.conf 가 선언한 그 노드의 GPU 수> / REPORTED=<gres.conf 가 실제로 가리키는 장치 수> (the placeholders are the number of GPUs slurm.conf declares for that node and the number of devices gres.conf actually points to)
Count the numbers in each of the two configuration files and compare. The numbers are printed in the log too.
Submit the corrected version
Submit /root/rca/slurm.conf and /root/rca/gres.conf with all four problems fixed. Make the problem node's CPUs match reality, remove the nonexistent node from the partition's node list, and make the GRES counts match. (munge is an operations problem, not a configuration file problem, so it is not dealt with in these files.)
You must fix all four problems. You can reuse the validation script you made in the previous lab.
Write the recovery procedure
Write /root/rca/recovery.sh. It must contain the following two kinds of commands, in order.
- The command that applies configuration changes (
scontrol reconfigure) - The command that brings the dropped nodes back (
scontrol update NodeName=... State=RESUME)
Applying the configuration and bringing nodes back are separate commands. The order matters too.
RCA report
Make /root/rca/report.txt with the following 6 lines.
CAUSE1=cpus-mismatch / CAUSE2=munge-<3번의 CAUSE 값> / CAUSE3=unknown-node / CAUSE4=gres-count / BAD_NODES=<1번 파일의 줄 수> / FIXED=yes (the placeholders are the CAUSE value from step 3 and the number of lines in the file from step 1)
Summarize the four causes under the prescribed keys. The values must be what you identified in the earlier steps.