Writing an sbatch Script
Goal
You write an sbatch script as required, set up an array job and a dependency chain, and create a validator that catches misplaced #SBATCH lines.
Why it matters
#SBATCH is parsed only until the first executable command appears. A directive written after that becomes an ordinary comment, and neither an error nor a warning appears. So you end up in a situation where you believed you requested a GPU but CUDA is not visible, and the user spends time suspecting the code. This error is hard for a person to catch by eye and easy to catch with a script.
And if you leave out the %N limit of an array job, 100 jobs fill the queue at once and block other users. On a shared cluster this is a social problem too.
Steps
- Create the
/root/jobsdirectory and create/root/jobs/train.sbatch. The first line is#!/bin/bash, followed by the three directives--job-name=resnet-train,--output=/root/jobs/logs/%x-%j.out, and--error=/root/jobs/logs/%x-%j.err. Also create the/root/jobs/logsdirectory. - Add the resource request to the same file.
--nodes=1,--ntasks=1,--cpus-per-task=8,--mem=64G - Add the time and the partition.
--time=04:00:00,--partition=gpu - Add the GPU request.
--gres=gpu:a100:2 - Create
/root/jobs/sweep.sbatch. It is an array job and must include--array=1-20%4and--output=/root/jobs/logs/%A_%a.out, and it must useSLURM_ARRAY_TASK_IDin the body. - Put the following in the body of
train.sbatch. It must start withset -euo pipefail, use theSLURM_CPUS_PER_TASKenvironment variable, and have a line that runs withsrun. - Create
/root/jobs/pipeline.sh. It must have two lines: one that receives the first job ID with--parsableand puts it in a variable, and one that submits the second job with--dependency=afterok:. - Write
/root/jobs/lint.sh. It takes the sbatch script path as its first argument, checks the following, and must exit with code 0 if there is no violation and with a nonzero value if there is.- Rule 1: The first line must start with
#! - Rule 2: Every
#SBATCHline must come before the first executable command (a line that is neither a comment nor blank) - Rule 3: The
--job-nameand--timedirectives must be present The grader runs it against both yourtrain.sbatch(which must pass) and a bad script with the directives deliberately placed after (which must fail).
- Rule 1: The first line must start with
Notes
- Output pattern substitutions:
%jjob ID,%xjob name,%Aarray parent ID,%aarray index. - Dependency example:
J1=$(sbatch --parsable prep.sh), then on the next linesbatch --dependency=afterok:$J1 train.sh - In this lab sbatch is not actually run. The contents of the script are graded.
- Common mistake 1: writing a directive without a space, like
#SBATCH--time. A space is needed after#SBATCH. - Common mistake 2: a step 8 validator that counts blank lines or comments as an 'executable command', and so judges a valid script as a failure.
The basic skeleton
Create the /root/jobs directory and create /root/jobs/train.sbatch. The first line is #!/bin/bash, followed by the three directives --job-name=resnet-train, --output=/root/jobs/logs/%x-%j.out, and --error=/root/jobs/logs/%x-%j.err. Also create the /root/jobs/logs directory.
Write the directives from the line after the shebang. The name and the output and error paths are the basics.
Resource request
Add the resource request to the same file. --nodes=1, --ntasks=1, --cpus-per-task=8, --mem=64G
Specify four things: nodes, tasks, CPUs per task, and memory.
Time and partition
Add the time and the partition. --time=04:00:00, --partition=gpu
The time format is hours:minutes:seconds. Use the partition you made in the previous lab.
GPU request
Add the GPU request. --gres=gpu:a100:2
The GRES format is name:type:count. Use the type you defined in the previous lab.
Array job
Create /root/jobs/sweep.sbatch. It is an array job and must include --array=1-20%4 and --output=/root/jobs/logs/%A_%a.out, and it must use SLURM_ARRAY_TASK_ID in the body.
After the range you can limit the number of concurrent runs with a percent sign. Put the array substitutions in the output pattern.
Environment variables and srun
Put the following in the body of train.sbatch. It must start with set -euo pipefail, use the SLURM_CPUS_PER_TASK environment variable, and have a line that runs with srun.
If you read the number of CPUs per task from the environment variable and use it, the resource request and the code automatically match.
Dependency chain
Create /root/jobs/pipeline.sh. It must have two lines: one that receives the first job ID with --parsable and puts it in a variable, and one that submits the second job with --dependency=afterok:.
There is an option that prints only the job ID. Make it continue only on success.
Script validator
Write /root/jobs/lint.sh. It takes the sbatch script path as its first argument, checks the following, and must exit with code 0 if there is no violation and with a nonzero value if there is.
- Rule 1: The first line must start with
#! - Rule 2: Every
#SBATCHline must come before the first executable command (a line that is neither a comment nor blank) - Rule 3: The
--job-nameand--timedirectives must be present The grader runs it against both yourtrain.sbatch(which must pass) and a bad script with the directives deliberately placed after (which must fail).
Directives after the first executable command are ignored. Catching that is the purpose of this validator.