GPU troubleshooting
My job won't start#
This is Why is my job pending? — GPU jobs use the same scheduler and the same reasons apply. A Resources reason specifically means Slurm is waiting for a GPU to actually become free; since only 3 nodes carry GPUs, that wait can be longer than for cpuqueue, but this page won't promise a wait time — see Partitions, QoS, and fairshare for how priority works.
My --gres request was rejected at submission#
If sbatch refuses the job immediately rather than queuing it, the request itself is invalid — most often more GPUs than exist on any single node. Verified maximum is 4 GPUs on one node (gpu:a100:4); requesting more than that on --nodes=1 cannot be satisfied. Check --gres=gpu:N against that ceiling, and check for a typo in the GRES name — the configured type is a100.
My job started, but my program says no GPU is available#
Work through these in order:
- Check
--greswas actually in your submission script. A job submitted without--gresruns fine on agpuqueuenode with zero GPUs allocated — Slurm doesn't require a GPU request just because you picked the GPU partition. - Check
$CUDA_VISIBLE_DEVICESwas set inside the job. Slurm sets this automatically when a GPU GRES is allocated (see Submit a GPU job). If it's empty, the allocation itself is the problem — check step 1. - Check your software actually loaded a GPU-capable build. A CPU-only install of a framework that also has a GPU build will run — just without using the GPU, and usually without an error. Confirm you loaded the right module version.
- Check the framework can see the device from inside Python/your runtime, not just that the environment variable is set — a driver or library-version mismatch between your software build and what's installed can cause the device to be invisible to the application even though Slurm allocated it correctly.
Note Do not use --mem to fix this — a device-visibility problem is not a memory problem, and increasing system RAM does nothing for it.
My job failed with an out-of-memory error — which kind?#
A "CUDA out of memory" error and a Slurm OUT_OF_MEMORY state are not the same failure — see the comparison below.
This is the distinction that matters most, and the two are unrelated:
| System memory (RAM) OOM | GPU memory (VRAM) OOM | |
|---|---|---|
| What ran out | The --mem/--mem-per-cpu you requested | The GPU's own memory — not requested through Slurm at all |
| How you'll see it | Slurm state OUT_OF_MEMORY in squeue/sacct | An error from your application/CUDA runtime (commonly something like "CUDA out of memory") in your job's output, not a Slurm state |
| What to check | sacct -j <jobid> --format=State,ReqMem,MaxRSS — see Why did my job fail? | Your program's own error output — Slurm has no visibility into VRAM usage |
| What actually fixes it | Increase --mem/--mem-per-cpu and resubmit | Reduce what your program asks the GPU to hold — smaller batch size, smaller model, gradient accumulation. Increasing --mem does nothing here. |
Important A CUDA/GPU out-of-memory error is an application-level failure, not a Slurm resource-limit failure. Slurm does not track or enforce GPU memory on this cluster, so there is no Slurm-side setting to raise. The fix lives in your job's own parameters (batch size, model size, precision), not in your #SBATCH directives.
My job ran, but seems to be using the CPU instead of the GPU#
Usually one of the causes under "no GPU is available" above, just discovered after the fact rather than from an explicit error — many frameworks silently fall back to CPU rather than failing loudly. Check your output logs for any framework-level warning about device selection, and verify with gpustat (see Submit a GPU job) during a short test run that the GPU is actually being used.
Driver, CUDA, or library version mismatches#
If your software reports a version-incompatibility error (a CUDA runtime that doesn't match the driver, or a library built against a different CUDA version than what's loaded), that's a software environment problem, not a scheduling one. Confirm exactly which module version you loaded, and check whether a different version of the same tool is available:
module avail <name>If nothing available fits, this is a Request software situation rather than something to work around.
Getting help#
Raise it through Getting help when you've worked through the relevant section above and the job still doesn't behave as expected. Include the job ID, your submission script, and — for a GPU-visibility or version problem — the exact error text from your program's output.
