Submitting jobs

Work on Mjolnir runs through Slurm. You describe the work in a script, submit it, and Slurm runs it on a compute node when suitable resources are free.

How a job flows#

login node          scheduler              compute node
-----------         ---------              ------------
write script   →    sbatch      →  PENDING  →  RUNNING  →  output file

You stay on the login node. Slurm does the placing. Your job's output is written to a file you name, which appears in the directory you submitted from.

A complete job script#

bash
#!/bin/bash
#SBATCH --job-name=example
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=02:00:00
#SBATCH --output=example-%j.out
#SBATCH --error=example-%j.err

echo "Running on $(hostname)"
# your commands here

Submit it:

bash
sbatch example.sh

Slurm responds with Submitted batch job <jobid>. Everything after the #SBATCH block is an ordinary shell script that runs on the compute node.

Important --qos is mandatory. Jobs submitted without a QoS are rejected. Scripts written before March 2025 predate this requirement and will need the line added.

The directives that matter#

DirectiveWhat it doesNotes
--job-nameLabel shown in the queueMakes squeue output readable
--partitionWhich set of machinescpuqueue is the default
--qosLimits and priorityRequired
--nodes=1Keep the job on one nodeRecommended unless your software runs across nodes
--cpus-per-taskCPU coresMust match what your program actually uses
--mem-per-cpuMemory per coreOr use --mem for a total
--timeMaximum runtimeThe job is killed at this limit
--output / --errorWhere output goes%j inserts the job ID

For the current partitions, QoS names, and their exact limits, see Partitions, QoS, and fairshare.

Requesting resources sensibly#

Two rules cover most of it:

Ask for what you will use. If you run a tool with 4 threads, request 4 CPUs. Requesting 32 does not make it faster — it makes the job wait longer for a larger free slot, and it counts against your fairshare as though you had used them.

Give a realistic --time. The job is terminated when the limit is reached, so too short loses your work. But a shorter, honest limit is easier for the scheduler to fit into a gap, so it often starts sooner.

Note Usage is charged on what you request, not what you consume. Over-requesting quietly lowers the priority of your future jobs.

A real job: loading the software it needs#

The example above runs shell commands. A research job also needs its software, and it has to load it itself:

bash
#!/bin/bash
#SBATCH --job-name=align
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=02:00:00
#SBATCH --output=align-%j.out
#SBATCH --error=align-%j.err

module purge
module load samtools/1.21

samtools --version
# your real command here, e.g.:
# samtools sort -@ $SLURM_CPUS_PER_TASK -o sorted.bam input.bam

Important A batch job does not inherit the modules you loaded in your terminal. It starts in its own environment on another machine, so a script that works interactively can still fail with "command not found". Load what you need inside the script.

Starting with module purge makes the job independent of whatever happened to be loaded when you submitted, and pinning the version (samtools/1.21) keeps the job reproducible when defaults change. Find and load software covers finding module names.

Matching CPUs to your program#

Requesting cores does not make software use them. Most tools need to be told:

bash
#SBATCH --cpus-per-task=8

my_tool --threads $SLURM_CPUS_PER_TASK input.fa

Using $SLURM_CPUS_PER_TASK keeps the request and the program in step, so changing the directive changes both.

Useful variables inside a job#

VariableContains
$SLURM_JOB_IDThe job ID
$SLURM_CPUS_PER_TASKCPUs allocated per task
$SLURM_JOB_NODELISTNode(s) the job is running on
$SLURM_SUBMIT_DIRDirectory the job was submitted from

After submitting#