Job arrays

When you need to run the same analysis over 50 samples, submitting 50 scripts is the wrong answer. A job array submits one script that Slurm runs many times, each with a different index.

Each task is scheduled independently, so they run in parallel when resources allow, and one task failing does not stop the others.

The idea#

one script  +  --array=1-50  →  50 independent tasks
                                each with its own $SLURM_ARRAY_TASK_ID

Your script uses that index to work out which input it should process.

A working example#

Say you have a file samples.txt with one sample name per line:

bash
#!/bin/bash
#SBATCH --job-name=per-sample
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=2
#SBATCH --mem-per-cpu=2G
#SBATCH --time=01:00:00
#SBATCH --array=1-50
#SBATCH --output=per-sample-%A_%a.out
#SBATCH --error=per-sample-%A_%a.err

module purge
module load samtools/1.21

SAMPLE=$(sed -n "${SLURM_ARRAY_TASK_ID}p" samples.txt)
echo "Task $SLURM_ARRAY_TASK_ID processing $SAMPLE on $(hostname)"

# your per-sample command here

Submit it exactly like any other job:

bash
sbatch per-sample.sh

Output naming#

Use both placeholders so tasks do not overwrite each other:

PlaceholderExpands to
%AThe array job ID (shared by all tasks)
%aThe task index
%jThe individual task's own job ID

per-sample-%A_%a.out gives you one clearly-named file per task.

Index ranges#

bash
#SBATCH --array=1-50        # 1 to 50
#SBATCH --array=0-9         # 0 to 9
#SBATCH --array=1,5,9       # only these
#SBATCH --array=1-100:2     # every second index

Pick the range that matches how your script maps an index to an input — off-by-one errors here are the most common array mistake. Test with a small range such as --array=1-2 before submitting the full set.

Limiting how many run at once#

Append % and a number to cap concurrent tasks:

bash
#SBATCH --array=1-100%10

This runs 100 tasks but never more than 10 at a time.

Tip Throttling is worth using when tasks are I/O-heavy, read the same shared files, or would otherwise flood the queue. It also makes a mistake cheaper — you find out something is wrong after 10 failures rather than 100.

Size limits#

The maximum array index on Mjolnir is 10000.

Warning Just because a large array is allowed does not make it a good idea. Thousands of very short tasks create more scheduling overhead than useful work — if each task takes seconds, group several inputs per task instead. And remember every task counts against your fairshare, so a huge array affects the priority of your later jobs.

Monitoring an array#

bash
squeue -u $USER

Tasks appear as <arrayjobid>_<index>. Pending tasks may be shown as a compressed range such as 1234567_[11-50].

After they finish:

bash
sacct -j <arrayjobid> --format=JobID,State,ExitCode,Elapsed,MaxRSS

This lists every task, so you can see which ones failed rather than assuming the whole array did.

Note If tasks sit pending with reason JobArrayTaskLimit, that is your own % throttle doing its job, not a problem. See Why is my job pending?.

Cancelling#

bash
scancel 1234567          # the whole array
scancel 1234567_7        # just task 7

See Cancelling jobs for what to expect afterward.