Partitions, QoS, and fairshare

Every job on Mjolnir needs two things: a partition (which machines it may run on) and a QoS (what limits and priority it runs under). This page is the reference for both, plus how Slurm decides whose job runs next.

Important Every job must specify a QoS. Jobs submitted without --qos are rejected. If you have scripts written before March 2025, they need updating.

Choosing a partition#

PartitionUse it forNodesMax timeDefault timeDefault memory/CPUMaximum memory
cpuqueueGeneral compute. The default.2014 days1 hour2 GBNode-dependent
gpuqueueJobs that need a GPU314 days1 hour2 GBNode-dependent
lazyqueueLow-priority work that can be interrupted2014 days1 hour2 GBNode-dependent
filetransferMoving large amounts of data144 days12 hours1 GB20 GB per job
teachingCourses and teaching sessions11 day10 hours2 GBNode-dependent

cpuqueue is the default: if you do not specify --partition, your job lands there. It includes the GPU nodes, because those machines have CPUs available too — but you will not get a GPU allocated unless you ask for one on gpuqueue.

filetransfer and teaching are restricted. filetransfer requires the mjolnir account and the filetransfer QoS; teaching requires the teaching account. If you are not in the relevant account, your job will be rejected rather than queued.

How much memory can I request?#

"Default memory/CPU" is only what you get if you don't ask for anything else — it is not a ceiling. On cpuqueue, gpuqueue, lazyqueue, and teaching, nothing in Slurm caps how much memory a job can request beyond what a real node actually has: there is no configured maximum, so the practical limit is node-dependent. Our nodes currently range from roughly 450 GB on the smallest to roughly 2 TB on the largest, with most nodes around 1 TB.

Note --mem and --mem-per-cpu both request memory per node, not a total pooled across every node in the job. A single-node job's request has to fit on one real node, however many CPUs it uses.

Requesting more memory narrows the set of nodes that can run your job — a request near the top of that range can only start on one of our largest nodes, so it may wait longer for one to be free. A request larger than any node actually has will never start at all. Neither is a reason to under-request: ask for what your application genuinely needs, and expect to wait longer as that number gets close to what our biggest nodes provide.

filetransfer is different: its QoS caps memory explicitly, at 20 GB per job (and 400 GB per user across all your filetransfer jobs) — see the QoS table below.

Choosing a QoS#

QoSPriorityMax CPUsMax timeNotes
normal100048 per job, 48 per nodePartition limit (14 days)The everyday default
teaching2000481 day (partition limit)Teaching account only
filetransfer20001 per job, 20 total per user4 daysAlso caps memory: 20 GB per job, 400 GB per user
lazy1No QoS limitPartition limit (14 days)Preemptible; see below

Warning The normal CPU limit is 48 per job and 48 per node — not 48 per user. There is no per-user CPU cap on normal. Older guidance stating "96 CPUs" or "48 CPUs per user" is out of date.

Because the limit applies per node as well as per job, a job asking for more than 48 CPUs will not start regardless of how it is spread. Add #SBATCH --nodes=1 to keep a job on one node and make that limit predictable.

Putting it together#

A standard compute job:

bash
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=2G
#SBATCH --time=04:00:00

A GPU job — --gres=gpu:N requests N GPUs on the node:

bash
#SBATCH --partition=gpuqueue
#SBATCH --qos=normal
#SBATCH --gres=gpu:1
#SBATCH --cpus-per-task=8
#SBATCH --time=08:00:00

See GPU computing for what's actually available and how to confirm your job received a GPU.

A data transfer:

bash
#SBATCH --partition=filetransfer
#SBATCH --qos=filetransfer
#SBATCH --cpus-per-task=1
#SBATCH --mem=20G
#SBATCH --time=1-00:00:00

The lazy queue#

lazyqueue runs work at the lowest possible priority on resources nobody else is using. It shares the same hardware as cpuqueue, so it is a way to get otherwise-idle capacity rather than a separate machine pool.

bash
#SBATCH --partition=lazyqueue
#SBATCH --qos=lazy

The trade-off is preemption: when a higher-priority job needs the resources, your job is stopped and requeued after a 5-minute grace period, then starts again from the beginning later. This cluster's default job-requeue setting already makes this work without any extra flag — see Lazyqueue: opportunistic, preemptible work for exactly why, and what to avoid.

Tip Use lazyqueue for work that tolerates being restarted from scratch and isn't urgent. Do not use it for anything with a deadline.

Because a preempted job restarts from scratch, work that checkpoints its progress benefits most. A job that would lose a week of computation when preempted is a poor fit.

How priority is decided#

Mjolnir uses Slurm's multifactor priority. When resources free up, the queued job with the highest score starts first. The factors and their weights:

FactorWeightMeaning
Fairshare10000Your recent usage relative to your share. Dominant factor.
Age2000How long the job has been waiting.
QoS1000The priority of the QoS you selected.
Job size1000Larger allocations score slightly higher.
Partition100Small per-partition adjustment.

Fairshare outweighs everything else combined. It falls as you consume resources and recovers as you idle, with a half-life of 14 days — so heavy use last week still affects your priority this week, and a quiet fortnight restores roughly half of it.

The scheduler also backfills: a small, short job can start ahead of a large queued one if it fits in a gap without delaying anything. Asking for a realistic --time makes your job easier to backfill, and asking for far more CPUs, memory or time than you need costs you priority.

Note Your usage is charged on what you request, not what you use. A job that reserves 48 CPUs and uses 4 is billed for 48 against your fairshare.

You can check your own fairshare standing with sshare -U.

Checking the current state yourself#

The values on this page are read from the live cluster, but the cluster is the authority. To check for yourself:

bash
sinfo -s                              # partitions and node states
scontrol show partition cpuqueue      # full detail for one partition
sinfo -N -p cpuqueue -o "%N %m %c"    # memory and CPUs per node
sacctmgr show qos format=Name,Priority,MaxWall,MaxTRESPerJob,MaxTRESPerUser
sshare -U                             # your fairshare