TT Lab
Get started
Learn Learning paths Courses

HPC and Slurm

How to Write an sbatch Script

Continue in TT Lab

In one line

The #SBATCH lines of an sbatch script look like comments, but they are a resource contract. And they are read only if they come before the first executable command.

Why this was needed

This is the most common beginner mistake.

#!/bin/bash
echo "starting"
#SBATCH --gres=gpu:2      # <- 무시된다!
python train.py

#SBATCH is parsed only until the first executable command appears. Anything after that is an ordinary comment. Neither an error nor a warning appears. The job is submitted without a GPU, and the user wanders for a long time wondering why CUDA is not visible.

How it works

The basic skeleton

#!/bin/bash
#SBATCH --job-name=resnet-train
#SBATCH --output=/home/me/logs/%x-%j.out
#SBATCH --error=/home/me/logs/%x-%j.err
#SBATCH --partition=gpu
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=64G
#SBATCH --gres=gpu:a100:2
#SBATCH --time=04:00:00

set -euo pipefail
echo "job $SLURM_JOB_ID on $SLURMD_NODENAME"
srun python train.py --epochs 50

It is good to know the substitution patterns for output file names.

Pattern Value
%j Job ID
%x Job name
%A The parent ID of an array job
%a The array index
%N The name of the first node

The directory of the output path must exist beforehand. If it does not, the job fails as soon as it starts, and since it cannot even create the file to hold the reason for the failure, it is hard to find the cause.

Useful environment variables

The values Slurm puts in inside a job.

SLURM_JOB_ID          작업 ID
SLURM_JOB_NAME        작업 이름
SLURM_JOB_NODELIST    할당된 노드 목록
SLURMD_NODENAME       현재 실행 중인 노드
SLURM_CPUS_PER_TASK   태스크당 CPU  <- DataLoader num_workers 에 쓰면 좋다
SLURM_ARRAY_TASK_ID   배열 인덱스
SLURM_NTASKS          태스크 총 개수

If you write it like num_workers=int(os.environ.get("SLURM_CPUS_PER_TASK", 4)), the resource request and the code automatically match.

The role of srun

If you use srun inside an sbatch script, a job step is created. It is needed to run in parallel across several nodes or tasks; for a single process you can do without it, but having it makes resource accounting accurate.

srun --ntasks=4 python ddp_train.py     # 4개 프로세스로

Array jobs

You use them to run the same script many times with only the parameters changed.

#SBATCH --array=1-100%10

1-100 is the index range, and %10 is the maximum number to run at the same time. Without this limit, 100 enter the queue at once and block other users.

python sweep.py --config "configs/exp${SLURM_ARRAY_TASK_ID}.yaml"

Dependencies

JOB1=$(sbatch --parsable prep.sh)
sbatch --dependency=afterok:$JOB1 train.sh
Condition Meaning
afterok:ID After that job ends with success
afterany:ID After it ends, whether success or failure
afternotok:ID After it ends in failure (for cleanup jobs)
singleton When none of my jobs with the same name exist

--parsable prints only the job ID, which makes it easy to put into a variable.

Checking status

squeue -u $USER
squeue -j 12345 -o '%.10i %.20j %.8T %.10M %.6D %R'
scontrol show job 12345
sacct -j 12345 --format=JobID,JobName,State,Elapsed,MaxRSS,ReqTRES
scancel 12345

The last column of squeue (%R) is the reason for waiting. Resources (waiting for resources), Priority (waiting on priority), Dependency (waiting on a dependency), QOSMaxJobsPerUserLimit (a limit applies), and so on appear. Half of the reasons a job is not running are written in this one column.

What it looks like in the field

Running with the default because --mem was not written. If the cluster default is small, the job dies of OOM, and if it is large, other jobs cannot get in. It is good to make a habit of stating it explicitly.

The culture of writing --time as the maximum. If everyone does it, backfill is crippled and overall throughput drops. Writing close to the actual time and leaving a margin of about 20% benefits everyone.

What you will do in the next lab

You write an sbatch script as required, and set up even an array job and a dependency chain. Finally, you create yourself a validation script that catches misplaced #SBATCH lines.