Submit a GPU job
A minimal working example#
#!/bin/bash
#SBATCH --job-name=gpu-example
#SBATCH --partition=gpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --gres=gpu:1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G
#SBATCH --time=00:10:00
#SBATCH --output=gpu-example-%j.out
module purge
gpustatEvery value here is deliberately small — this is a template to build from, not a recommendation for real workloads. --cpus-per-task=4 and --mem-per-cpu=2G are ordinary starting points, not a required ratio to --gres=gpu:1: there is no configured or documented CPU-per-GPU rule on this cluster, so size them to what your actual program needs.
--qos=normal is mandatory, exactly as everywhere else — there is no separate GPU QoS.
Confirm you actually got a GPU#
nvidia-smi is live-verified as available and working inside a GPU allocation — a single-GPU probe job on gpu01 ran it successfully and used it to identify the allocated device:
nvidia-smiUse it to confirm which GPU you were given and to check its current device-memory and utilization status — it doesn't track historical usage, and there's no GPU on the login node for it to report on, so it's only useful inside a running allocation. The probe observed one NVIDIA A100-PCIE-40GB with 40960 MiB of device memory — a real example from this cluster, not a claim that every Mjolnir A100 reports identical capacity.
A gpustat module is also available — verified present in the module system — documented upstream as a more readable wrapper around that same NVIDIA tooling:
module load gpustat
gpustatNote gpustat's own output inside a Mjolnir GPU allocation hasn't been independently observed — nvidia-smi above is the live-verified option. If gpustat's output doesn't match what you expect, treat that as worth asking about rather than assuming you've misread it.
Inside a GPU allocation, Slurm sets CUDA_VISIBLE_DEVICES to the specific device(s) it assigned you — live-verified: the same probe observed CUDA_VISIBLE_DEVICES=0 for its one-GPU allocation. The exact index depends on which device Slurm happens to give you; what matters is the mechanism — Slurm, not your application, controls which device(s) are exposed to the job. Don't override it: doing so tells your application to look at a different device than the one the scheduler actually gave you, which either fails outright or, worse, quietly competes with someone else's job.
echo "Assigned GPU(s): $CUDA_VISIBLE_DEVICES"Load your GPU-capable software#
Some packaged GPU software carries its own required runtime with it. For example, the pytorch/2.2.2 module is a self-contained environment — inspected directly, it needs no separate module load cuda to work:
module purge
module load pytorch/2.2.2
python your_script.pyDon't assume every GPU-capable module behaves the same way — check what a given module actually needs with module show <name> before assuming it does or doesn't need cuda loaded alongside it.
The cuda module (several versions are available) is for compiling your own CUDA code directly — if you're using an existing packaged framework rather than writing CUDA code yourself, check first with module show whether it's actually needed.
Tip Test with a tiny job first — a few minutes, one GPU — before submitting anything long. It's much cheaper to discover a broken environment or an unexpected module conflict in a 5-minute job than a 12-hour one.
Full pattern with a real framework#
#!/bin/bash
#SBATCH --job-name=train
#SBATCH --partition=gpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --gres=gpu:1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=4G
#SBATCH --time=04:00:00
#SBATCH --output=train-%j.out
module purge
module load pytorch/2.2.2
python train.pyAdjust --cpus-per-task and --mem-per-cpu to your actual data-loading needs — more workers generally means more CPUs, but this is not a fixed rule.
After it runs: what did you actually get?#
sacct -j <jobid> --format=JobID,State,ExitCode,Elapsed,AllocTRESAllocTRES includes gres/gpu=N, confirming exactly how many GPUs Slurm allocated to the job — useful for confirming a job that behaved oddly really did get the GPU count you asked for.
For everything else about monitoring and interpreting job state, see Monitoring jobs — GPU jobs behave like any other job for squeue, scontrol show job, and the rest of sacct.
Multiple GPUs#
#SBATCH --gres=gpu:4This requests 4 GPUs on one node. Requesting more GPUs does not make a single-GPU program faster — your application has to be written to actually use more than one device (data-parallel or model-parallel training, for example). Check whether your framework/script supports multi-GPU before requesting more than one; otherwise you're just holding hardware idle.
Note Multi-node GPU jobs (spanning more than one of the three GPU nodes) aren't covered here — there's no established Mjolnir workflow or evidence for that pattern. If you need it, ask through Getting help rather than assuming it works the way it might elsewhere.
