What Happens When You Share GPUs by Word of Mouth
In one line
A batch scheduler is a system that decides "who uses what, and when" by policy rather than by agreement among people.
Why this was needed
When there is one GPU server, it goes like this. "Are you using it right now?" "No, go ahead." It works fine.
Once there are two servers and five users, a reservation thread appears in Slack. Once there are three servers and ten users, this kind of thing happens.
- Someone checked with
nvidia-smiand launched, but in the meantime someone else launched first, and both die of OOM - A job that was to run overnight finished at 3 a.m., but the GPU sat idle until 9 a.m.
- There is an urgent experiment, and nobody knows when the person ahead's job will finish
- Someone grabbed all 8 cards by mistake and forgot about them
What these problems have in common is that resource allocation is not managed as state. A scheduler turns it into a queue and policy.
How it works
The components of Slurm
| Component | Role | Where it runs |
|---|---|---|
slurmctld |
The central controller. Queue management and scheduling decisions | 1 management node (2 for HA) |
slurmd |
The agent on each compute node. Runs jobs | All compute nodes |
slurmdbd |
The accounting and usage database (optional) | Management node |
munge |
Authentication between nodes | All nodes |
munge causes problems most often in this picture. Every node in the cluster must have the same munge key, and the clocks must also agree (the default tolerance is a few minutes). If the keys differ or the clocks drift, Munge decode failed appears and the nodes cannot communicate.
ls -l /etc/munge/munge.key # 0400, munge:munge 여야 한다
If the permissions are even slightly loose, munged does not come up at all.
Basic concepts of scheduling
- Node — the unit of compute resources. It has CPUs, memory, and GRES (GPUs and so on).
- Partition — a group of nodes and also a unit of policy. It sets the maximum run time, priority, and access permissions. It corresponds to the "queue" of other schedulers.
- Job — a resource request submitted by a user plus what to run.
- Job step (Step) — a unit run inside a job with
srun.
Backfill
The default scheduler places jobs in priority order. But while a large job waits for resources, nodes can sit idle. Backfill fills in smaller jobs behind it first, "within the range that does not delay the start time of the job ahead".
For backfill to work well, users must write --time honestly. If everyone writes the maximum, the scheduler cannot compute the gaps. That is why mature clusters turn "write it short and it runs sooner" into an incentive by policy.
What a resource request means
sbatch --nodes=1 --ntasks=1 --cpus-per-task=8 --mem=64G --gres=gpu:a100:2 --time=04:00:00 train.sh
--nodes— the number of nodes--ntasks— the number of tasks (processes) to run. Corresponds to the number of MPI ranks--cpus-per-task— CPU cores per task. It is directly tied to the number of PyTorch DataLoader workers--mem/--mem-per-cpu— memory. Use only one of the two--gres=gpu:<타입>:<개수>— GPUs (the placeholders are the type and the count)
You cannot use resources you did not request. It is enforced by cgroup, so if you get --gres=gpu:1 and the code tries to use 2 cards, the second card is not visible at all.
How to read why the queue is not running
Once you have attached a batch scheduler, the thing you hear most often is "my job isn't running". Slurm tells you the reason as a state string, so if you just know the table, most cases resolve themselves.
squeue -u $USER -o "%.10i %.9P %.20j %.8T %.10M %R"
scontrol show job <작업번호> | grep -E 'JobState|Reason|NodeList'
| Reason | Meaning | The usual fix |
|---|---|---|
Resources |
The requested resources are not free right now | Wait. A smaller request makes it faster |
Priority |
A higher-priority job is ahead | A matter of the fair-share policy |
QOSMaxJobsPerUserLimit |
The concurrent-run cap of that QoS | Bundle them into an array job |
AssocGrpCPUMinutesLimit |
The account's allocation has been used up | Ask the administrator |
ReqNodeNotAvail |
That node is reserved or down | Release the node specification |
PartitionTimeLimit |
The requested time is longer than the partition limit | Reduce --time |
The bigger you make the request, the longer you wait. If you request 4 nodes for 8 hours, you wait until a free window that size appears. The backfill scheduler fits short, small jobs into the gaps, so if you write --time honestly (generous but not excessive), the job starts much sooner. The habit of writing the maximum limit delays yourself.
If a node is drain, the reason is written.
sinfo -R # drain 사유 목록
scontrol show node <노드> | grep -E 'State|Reason'
Most are disk shortage, GPU errors, or health check failures. If you bring the node back without clearing the reason, it drops out again at the next check.
When a job died and you do not know why, look at the accounting records. When it vanishes without leaving anything in standard output, sacct tells you OUT_OF_MEMORY or TIMEOUT along with the exit code.
sacct -j <작업번호> --format=JobID,State,ExitCode,MaxRSS,Elapsed,ReqMem
If MaxRSS is close to ReqMem, memory is the cause.
What it looks like in the field
It makes sense in a homelab too. Even on a server with two GPUs, if you have to run several experiments in sequence, a scheduler is better. If you put them in the queue overnight, they run one after another automatically. And that experience carries over directly to a large cluster.
Nodes dropping into DRAIN. When a hardware error or a GRES mismatch is detected, Slurm removes that node automatically. After fixing the cause, you must bring it back with scontrol update NodeName=... State=RESUME. It does not come back automatically — if you do not know this, a single node sits idle for days.
What to look for in the next check
In the quiz that follows, you first check which scheduling decision each of partitions, backfill, and resource requests produces. After passing those criteria, from the next module on you practice writing slurm.conf, defining GPU GRES, and diagnosing sbatch scripts and broken configurations, in order.