Submitting jobs
Work on Mjolnir runs through Slurm. You describe the work in a script, submit it, and Slurm runs it on a compute node when suitable resources are free.
How a job flows#
login node scheduler compute node
----------- --------- ------------
write script → sbatch → PENDING → RUNNING → output fileYou stay on the login node. Slurm does the placing. Your job's output is written to a file you name, which appears in the directory you submitted from.
A complete job script#
#!/bin/bash
#SBATCH --job-name=example
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=02:00:00
#SBATCH --output=example-%j.out
#SBATCH --error=example-%j.err
echo "Running on $(hostname)"
# your commands hereSubmit it:
sbatch example.shSlurm responds with Submitted batch job <jobid>. Everything after the #SBATCH block is an ordinary shell script that runs on the compute node.
Important --qos is mandatory. Jobs submitted without a QoS are rejected. Scripts written before March 2025 predate this requirement and will need the line added.
The directives that matter#
| Directive | What it does | Notes |
|---|---|---|
--job-name | Label shown in the queue | Makes squeue output readable |
--partition | Which set of machines | cpuqueue is the default |
--qos | Limits and priority | Required |
--nodes=1 | Keep the job on one node | Recommended unless your software runs across nodes |
--cpus-per-task | CPU cores | Must match what your program actually uses |
--mem-per-cpu | Memory per core | Or use --mem for a total |
--time | Maximum runtime | The job is killed at this limit |
--output / --error | Where output goes | %j inserts the job ID |
For the current partitions, QoS names, and their exact limits, see Partitions, QoS, and fairshare.
Requesting resources sensibly#
Two rules cover most of it:
Ask for what you will use. If you run a tool with 4 threads, request 4 CPUs. Requesting 32 does not make it faster — it makes the job wait longer for a larger free slot, and it counts against your fairshare as though you had used them.
Give a realistic --time. The job is terminated when the limit is reached, so too short loses your work. But a shorter, honest limit is easier for the scheduler to fit into a gap, so it often starts sooner.
Note Usage is charged on what you request, not what you consume. Over-requesting quietly lowers the priority of your future jobs.
A real job: loading the software it needs#
The example above runs shell commands. A research job also needs its software, and it has to load it itself:
#!/bin/bash
#SBATCH --job-name=align
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=02:00:00
#SBATCH --output=align-%j.out
#SBATCH --error=align-%j.err
module purge
module load samtools/1.21
samtools --version
# your real command here, e.g.:
# samtools sort -@ $SLURM_CPUS_PER_TASK -o sorted.bam input.bamImportant A batch job does not inherit the modules you loaded in your terminal. It starts in its own environment on another machine, so a script that works interactively can still fail with "command not found". Load what you need inside the script.
Starting with module purge makes the job independent of whatever happened to be loaded when you submitted, and pinning the version (samtools/1.21) keeps the job reproducible when defaults change. Find and load software covers finding module names.
Matching CPUs to your program#
Requesting cores does not make software use them. Most tools need to be told:
#SBATCH --cpus-per-task=8
my_tool --threads $SLURM_CPUS_PER_TASK input.faUsing $SLURM_CPUS_PER_TASK keeps the request and the program in step, so changing the directive changes both.
Useful variables inside a job#
| Variable | Contains |
|---|---|
$SLURM_JOB_ID | The job ID |
$SLURM_CPUS_PER_TASK | CPUs allocated per task |
$SLURM_JOB_NODELIST | Node(s) the job is running on |
$SLURM_SUBMIT_DIR | Directory the job was submitted from |
After submitting#
- Watch it — Monitoring jobs
- It is not starting — Why is my job pending?
- Cancel it —
scancel <jobid>
Related#
- First 15 minutes on Mjolnir — the guided version of this page
- Where should my files go? — submit from a sensible location
- Check job efficiency — check the request against what the job used
- Getting help — when a job fails for reasons you cannot see
