TT Lab
Get started
Learn Learning paths Courses

HPC and Slurm

What Happens When You Share GPUs by Word of Mouth

Continue in TT Lab

In one line

A batch scheduler is a system that decides "who uses what, and when" by policy rather than by agreement among people.

Why this was needed

When there is one GPU server, it goes like this. "Are you using it right now?" "No, go ahead." It works fine.

Once there are two servers and five users, a reservation thread appears in Slack. Once there are three servers and ten users, this kind of thing happens.

What these problems have in common is that resource allocation is not managed as state. A scheduler turns it into a queue and policy.

How it works

The components of Slurm

Component Role Where it runs
slurmctld The central controller. Queue management and scheduling decisions 1 management node (2 for HA)
slurmd The agent on each compute node. Runs jobs All compute nodes
slurmdbd The accounting and usage database (optional) Management node
munge Authentication between nodes All nodes

munge causes problems most often in this picture. Every node in the cluster must have the same munge key, and the clocks must also agree (the default tolerance is a few minutes). If the keys differ or the clocks drift, Munge decode failed appears and the nodes cannot communicate.

ls -l /etc/munge/munge.key      # 0400, munge:munge 여야 한다

If the permissions are even slightly loose, munged does not come up at all.

Basic concepts of scheduling

Backfill

The default scheduler places jobs in priority order. But while a large job waits for resources, nodes can sit idle. Backfill fills in smaller jobs behind it first, "within the range that does not delay the start time of the job ahead".

For backfill to work well, users must write --time honestly. If everyone writes the maximum, the scheduler cannot compute the gaps. That is why mature clusters turn "write it short and it runs sooner" into an incentive by policy.

What a resource request means

sbatch --nodes=1 --ntasks=1 --cpus-per-task=8 --mem=64G --gres=gpu:a100:2 --time=04:00:00 train.sh

You cannot use resources you did not request. It is enforced by cgroup, so if you get --gres=gpu:1 and the code tries to use 2 cards, the second card is not visible at all.

How to read why the queue is not running

Once you have attached a batch scheduler, the thing you hear most often is "my job isn't running". Slurm tells you the reason as a state string, so if you just know the table, most cases resolve themselves.

squeue -u $USER -o "%.10i %.9P %.20j %.8T %.10M %R"
scontrol show job <작업번호> | grep -E 'JobState|Reason|NodeList'
Reason Meaning The usual fix
Resources The requested resources are not free right now Wait. A smaller request makes it faster
Priority A higher-priority job is ahead A matter of the fair-share policy
QOSMaxJobsPerUserLimit The concurrent-run cap of that QoS Bundle them into an array job
AssocGrpCPUMinutesLimit The account's allocation has been used up Ask the administrator
ReqNodeNotAvail That node is reserved or down Release the node specification
PartitionTimeLimit The requested time is longer than the partition limit Reduce --time

The bigger you make the request, the longer you wait. If you request 4 nodes for 8 hours, you wait until a free window that size appears. The backfill scheduler fits short, small jobs into the gaps, so if you write --time honestly (generous but not excessive), the job starts much sooner. The habit of writing the maximum limit delays yourself.

If a node is drain, the reason is written.

sinfo -R          # drain 사유 목록
scontrol show node <노드> | grep -E 'State|Reason'

Most are disk shortage, GPU errors, or health check failures. If you bring the node back without clearing the reason, it drops out again at the next check.

When a job died and you do not know why, look at the accounting records. When it vanishes without leaving anything in standard output, sacct tells you OUT_OF_MEMORY or TIMEOUT along with the exit code.

sacct -j <작업번호> --format=JobID,State,ExitCode,MaxRSS,Elapsed,ReqMem

If MaxRSS is close to ReqMem, memory is the cause.

What it looks like in the field

It makes sense in a homelab too. Even on a server with two GPUs, if you have to run several experiments in sequence, a scheduler is better. If you put them in the queue overnight, they run one after another automatically. And that experience carries over directly to a large cluster.

Nodes dropping into DRAIN. When a hardware error or a GRES mismatch is detected, Slurm removes that node automatically. After fixing the cause, you must bring it back with scontrol update NodeName=... State=RESUME. It does not come back automatically — if you do not know this, a single node sits idle for days.

What to look for in the next check

In the quiz that follows, you first check which scheduling decision each of partitions, backfill, and resource requests produces. After passing those criteria, from the next module on you practice writing slurm.conf, defining GPU GRES, and diagnosing sbatch scripts and broken configurations, in order.