GPU computing

What Mjolnir provides#

Three nodes carry NVIDIA A100 GPUs, live-verified: gpuqueue has 3 nodes and 640 CPUs total, with 2 GPUs on one node and 4 GPUs each on the other two — up to 4 A100s available on a single node.

When to request a GPU#

Only for work that is actually written to use one — deep learning training and inference, and GPU-accelerated tools that say so explicitly. Requesting a GPU for a CPU-only program gets you nothing: the hardware sits idle while your code runs on the CPU exactly as it would on cpuqueue, except now you're competing for a much scarcer resource.

Note If you're not sure whether your tool uses the GPU, check its documentation before requesting one. "Runs faster on some machines" is not the same as "uses CUDA."

How GPU scheduling differs from CPU work#

Mechanically, very little. There is no separate GPU QoS — verified directly: gpuqueue accepts the same QoS as everywhere else (AllowQos=ALL), and sacctmgr show qos shows no GPU-specific limits. You still pick --partition=gpuqueue, still specify a mandatory --qos (normal for ordinary work), and the same fairshare and CPU limits from Partitions, QoS, and fairshare apply — including the 48-CPU cap under normal.

The one addition is the GPU request itself.

CPU and system RAM are still separate from the GPU#

A GPU request does not automatically come with any particular amount of CPU or system RAM — --cpus-per-task and --mem/--mem-per-cpu mean exactly what they mean for a CPU job, and you request them independently of --gres. There is no configured or documented ratio between GPU count and CPU count on Mjolnir — asking for 1 GPU and 8 CPUs, or 1 GPU and 2 CPUs, are both valid; size the CPU and memory request to what your data loading and preprocessing actually need, the same way you would for any job.

Important --mem controls system RAM, not GPU memory. The two are entirely separate resources — see GPU memory (VRAM) below.

How do I request a GPU?#

bash
#SBATCH --partition=gpuqueue
#SBATCH --qos=normal
#SBATCH --gres=gpu:1

--gres=gpu:N requests N GPUs on one node, N from 1 to 4 depending on which node you land on. Every GPU on Mjolnir is the same model (A100), so this untyped form is normally all you need — the type is still there if you ever want to be explicit (--gres=gpu:a100:1), but it makes no practical difference right now since there's only one type.

See Submit a GPU job for a complete, working script.

GPU memory (VRAM)#

GPU memory is not tracked or requestable through Slurm on this cluster — there is no --mem equivalent for VRAM. It's entirely up to your application to fit within whatever the GPU you land on provides, and to fail informatively if it doesn't. See GPU troubleshooting for what a VRAM-exhaustion failure looks like and how it differs from an ordinary system-memory failure.

How do I see whether my job received a GPU?#

Covered in Submit a GPU job, which walks through checking the allocation before and after your job runs.

Where to go next#

To do thisRead
Submit a working GPU jobSubmit a GPU job
Diagnose a GPU job that isn't workingGPU troubleshooting
See exact partition/QoS limitsPartitions, QoS, and fairshare