How to Write an sbatch Script
In one line
The #SBATCH lines of an sbatch script look like comments, but they are a resource contract. And they are read only if they come before the first executable command.
Why this was needed
This is the most common beginner mistake.
#!/bin/bash
echo "starting"
#SBATCH --gres=gpu:2 # <- 무시된다!
python train.py
#SBATCH is parsed only until the first executable command appears. Anything after that is an ordinary comment. Neither an error nor a warning appears. The job is submitted without a GPU, and the user wanders for a long time wondering why CUDA is not visible.
How it works
The basic skeleton
#!/bin/bash
#SBATCH --job-name=resnet-train
#SBATCH --output=/home/me/logs/%x-%j.out
#SBATCH --error=/home/me/logs/%x-%j.err
#SBATCH --partition=gpu
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=64G
#SBATCH --gres=gpu:a100:2
#SBATCH --time=04:00:00
set -euo pipefail
echo "job $SLURM_JOB_ID on $SLURMD_NODENAME"
srun python train.py --epochs 50
It is good to know the substitution patterns for output file names.
| Pattern | Value |
|---|---|
%j |
Job ID |
%x |
Job name |
%A |
The parent ID of an array job |
%a |
The array index |
%N |
The name of the first node |
The directory of the output path must exist beforehand. If it does not, the job fails as soon as it starts, and since it cannot even create the file to hold the reason for the failure, it is hard to find the cause.
Useful environment variables
The values Slurm puts in inside a job.
SLURM_JOB_ID 작업 ID
SLURM_JOB_NAME 작업 이름
SLURM_JOB_NODELIST 할당된 노드 목록
SLURMD_NODENAME 현재 실행 중인 노드
SLURM_CPUS_PER_TASK 태스크당 CPU <- DataLoader num_workers 에 쓰면 좋다
SLURM_ARRAY_TASK_ID 배열 인덱스
SLURM_NTASKS 태스크 총 개수
If you write it like num_workers=int(os.environ.get("SLURM_CPUS_PER_TASK", 4)), the resource request and the code automatically match.
The role of srun
If you use srun inside an sbatch script, a job step is created. It is needed to run in parallel across several nodes or tasks; for a single process you can do without it, but having it makes resource accounting accurate.
srun --ntasks=4 python ddp_train.py # 4개 프로세스로
Array jobs
You use them to run the same script many times with only the parameters changed.
#SBATCH --array=1-100%10
1-100 is the index range, and %10 is the maximum number to run at the same time. Without this limit, 100 enter the queue at once and block other users.
python sweep.py --config "configs/exp${SLURM_ARRAY_TASK_ID}.yaml"
Dependencies
JOB1=$(sbatch --parsable prep.sh)
sbatch --dependency=afterok:$JOB1 train.sh
| Condition | Meaning |
|---|---|
afterok:ID |
After that job ends with success |
afterany:ID |
After it ends, whether success or failure |
afternotok:ID |
After it ends in failure (for cleanup jobs) |
singleton |
When none of my jobs with the same name exist |
--parsable prints only the job ID, which makes it easy to put into a variable.
Checking status
squeue -u $USER
squeue -j 12345 -o '%.10i %.20j %.8T %.10M %.6D %R'
scontrol show job 12345
sacct -j 12345 --format=JobID,JobName,State,Elapsed,MaxRSS,ReqTRES
scancel 12345
The last column of squeue (%R) is the reason for waiting. Resources (waiting for resources), Priority (waiting on priority), Dependency (waiting on a dependency), QOSMaxJobsPerUserLimit (a limit applies), and so on appear. Half of the reasons a job is not running are written in this one column.
What it looks like in the field
Running with the default because --mem was not written. If the cluster default is small, the job dies of OOM, and if it is large, other jobs cannot get in. It is good to make a habit of stating it explicitly.
The culture of writing --time as the maximum. If everyone does it, backfill is crippled and overall throughput drops. Writing close to the actual time and leaving a margin of about 20% benefits everyone.
What you will do in the next lab
You write an sbatch script as required, and set up even an array job and a dependency chain. Finally, you create yourself a validation script that catches misplaced #SBATCH lines.