Job arrays
When you need to run the same analysis over 50 samples, submitting 50 scripts is the wrong answer. A job array submits one script that Slurm runs many times, each with a different index.
Each task is scheduled independently, so they run in parallel when resources allow, and one task failing does not stop the others.
The idea#
one script + --array=1-50 → 50 independent tasks
each with its own $SLURM_ARRAY_TASK_IDYour script uses that index to work out which input it should process.
A working example#
Say you have a file samples.txt with one sample name per line:
#!/bin/bash
#SBATCH --job-name=per-sample
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=2
#SBATCH --mem-per-cpu=2G
#SBATCH --time=01:00:00
#SBATCH --array=1-50
#SBATCH --output=per-sample-%A_%a.out
#SBATCH --error=per-sample-%A_%a.err
module purge
module load samtools/1.21
SAMPLE=$(sed -n "${SLURM_ARRAY_TASK_ID}p" samples.txt)
echo "Task $SLURM_ARRAY_TASK_ID processing $SAMPLE on $(hostname)"
# your per-sample command hereSubmit it exactly like any other job:
sbatch per-sample.shOutput naming#
Use both placeholders so tasks do not overwrite each other:
| Placeholder | Expands to |
|---|---|
%A | The array job ID (shared by all tasks) |
%a | The task index |
%j | The individual task's own job ID |
per-sample-%A_%a.out gives you one clearly-named file per task.
Index ranges#
#SBATCH --array=1-50 # 1 to 50
#SBATCH --array=0-9 # 0 to 9
#SBATCH --array=1,5,9 # only these
#SBATCH --array=1-100:2 # every second indexPick the range that matches how your script maps an index to an input — off-by-one errors here are the most common array mistake. Test with a small range such as --array=1-2 before submitting the full set.
Limiting how many run at once#
Append % and a number to cap concurrent tasks:
#SBATCH --array=1-100%10This runs 100 tasks but never more than 10 at a time.
Tip Throttling is worth using when tasks are I/O-heavy, read the same shared files, or would otherwise flood the queue. It also makes a mistake cheaper — you find out something is wrong after 10 failures rather than 100.
Size limits#
The maximum array index on Mjolnir is 10000.
Warning Just because a large array is allowed does not make it a good idea. Thousands of very short tasks create more scheduling overhead than useful work — if each task takes seconds, group several inputs per task instead. And remember every task counts against your fairshare, so a huge array affects the priority of your later jobs.
Monitoring an array#
squeue -u $USERTasks appear as <arrayjobid>_<index>. Pending tasks may be shown as a compressed range such as 1234567_[11-50].
After they finish:
sacct -j <arrayjobid> --format=JobID,State,ExitCode,Elapsed,MaxRSSThis lists every task, so you can see which ones failed rather than assuming the whole array did.
Note If tasks sit pending with reason JobArrayTaskLimit, that is your own % throttle doing its job, not a problem. See Why is my job pending?.
Cancelling#
scancel 1234567 # the whole array
scancel 1234567_7 # just task 7See Cancelling jobs for what to expect afterward.
Related#
- Submitting jobs — the directives used above
- Monitoring jobs
- Why did my job fail? — when some tasks fail
- Partitions, QoS, and fairshare — limits and priority
