TT Lab
Get started
Learn Learning paths Courses

HPC and Slurm

Diagnosing a Broken Cluster

Continue in TT Lab

Goal

You receive a broken Slurm configuration and logs, identify each of the four causes, fix them, and leave a recovery procedure and an RCA report.

Why it matters

A characteristic of Slurm failures is that a single log line tells you the whole cause. The problem is that if you do not know the order for finding that one line, it takes days. The order is fixed — narrow the scope with sinfo, read the reason in slurmctld.log, and confirm that line in the configuration file.

It is especially important to tell the two kinds of munge error apart. Expired credential is a clock problem, so you look at NTP, and Invalid credential is a key problem, so you compare the key files. They are completely different responses.

The fixtures are in /opt/fixtures/slurm/broken/. They are four files: slurm.conf, gres.conf, slurmctld.log, and sinfo.txt.

Steps

  1. Create the /root/rca directory, and write, one per line, the distinct node names in a non-normal state (not idle/alloc/mix) from sinfo.txt in /root/rca/bad-nodes.txt.
  2. Find the node definition mismatch and write it in /root/rca/node.txt as two lines. NODE=<문제 노드 이름> / KEY=<불일치가 난 설정 키 이름> (the placeholders are the problem node name and the name of the configuration key that mismatched)
  3. Identify the munge authentication failure and write it in /root/rca/munge.txt as two lines. KIND=<expired|invalid> / CAUSE=<clock|key>
  4. Write, in one line in /root/rca/partition.txt, the node name that a partition references but that is not defined.
  5. Write the GRES count mismatch in /root/rca/gres.txt as two lines. CONFIGURED=<slurm.conf 가 선언한 그 노드의 GPU 수> / REPORTED=<gres.conf 가 실제로 가리키는 장치 수> (the placeholders are the number of GPUs slurm.conf declares for that node and the number of devices gres.conf actually points to)
  6. Submit /root/rca/slurm.conf and /root/rca/gres.conf with all four problems fixed. Make the problem node's CPUs match reality, remove the nonexistent node from the partition's node list, and make the GRES counts match. (munge is an operations problem, not a configuration file problem, so it is not dealt with in these files.)
  7. Write /root/rca/recovery.sh. It must contain the following two kinds of commands, in order.
    • The command that applies configuration changes (scontrol reconfigure)
    • The command that brings the dropped nodes back (scontrol update NodeName=... State=RESUME)
  8. Make /root/rca/report.txt with the following 6 lines. CAUSE1=cpus-mismatch / CAUSE2=munge-<3번의 CAUSE 값> / CAUSE3=unknown-node / CAUSE4=gres-count / BAD_NODES=<1번 파일의 줄 수> / FIXED=yes (the placeholders are the CAUSE value from step 3 and the number of lines in the file from step 1)

Notes

Summarize the symptoms

Create the /root/rca directory, and write, one per line, the distinct node names in a non-normal state (not idle/alloc/mix) from sinfo.txt in /root/rca/bad-nodes.txt.

Count the nodes in a non-normal state in the sinfo output. Write the state strings exactly.

Find the node definition mismatch

Find the node definition mismatch and write it in /root/rca/node.txt as two lines. NODE=<문제 노드 이름> / KEY=<불일치가 난 설정 키 이름> (the placeholders are the problem node name and the name of the configuration key that mismatched)

The log says which node dropped out and why. Confirm that node's line in the configuration file.

Identify the authentication failure

Identify the munge authentication failure and write it in /root/rca/munge.txt as two lines. KIND=<expired|invalid> / CAUSE=<clock|key>

There are two kinds of munge error. Look at the wording in the log as it is and decide which one it is.

Reference to an undefined node

Write, in one line in /root/rca/partition.txt, the node name that a partition references but that is not defined.

Compare the partition's node list with the nodes actually defined. You must expand the range notation.

GRES count mismatch

Write the GRES count mismatch in /root/rca/gres.txt as two lines. CONFIGURED=<slurm.conf 가 선언한 그 노드의 GPU 수> / REPORTED=<gres.conf 가 실제로 가리키는 장치 수> (the placeholders are the number of GPUs slurm.conf declares for that node and the number of devices gres.conf actually points to)

Count the numbers in each of the two configuration files and compare. The numbers are printed in the log too.

Submit the corrected version

Submit /root/rca/slurm.conf and /root/rca/gres.conf with all four problems fixed. Make the problem node's CPUs match reality, remove the nonexistent node from the partition's node list, and make the GRES counts match. (munge is an operations problem, not a configuration file problem, so it is not dealt with in these files.)

You must fix all four problems. You can reuse the validation script you made in the previous lab.

Write the recovery procedure

Write /root/rca/recovery.sh. It must contain the following two kinds of commands, in order.

Applying the configuration and bringing nodes back are separate commands. The order matters too.

RCA report

Make /root/rca/report.txt with the following 6 lines. CAUSE1=cpus-mismatch / CAUSE2=munge-<3번의 CAUSE 값> / CAUSE3=unknown-node / CAUSE4=gres-count / BAD_NODES=<1번 파일의 줄 수> / FIXED=yes (the placeholders are the CAUSE value from step 3 and the number of lines in the file from step 1)

Summarize the four causes under the prescribed keys. The values must be what you identified in the earlier steps.