GPU troubleshooting

My job won't start#

This is Why is my job pending? — GPU jobs use the same scheduler and the same reasons apply. A Resources reason specifically means Slurm is waiting for a GPU to actually become free; since only 3 nodes carry GPUs, that wait can be longer than for cpuqueue, but this page won't promise a wait time — see Partitions, QoS, and fairshare for how priority works.

My --gres request was rejected at submission#

If sbatch refuses the job immediately rather than queuing it, the request itself is invalid — most often more GPUs than exist on any single node. Verified maximum is 4 GPUs on one node (gpu:a100:4); requesting more than that on --nodes=1 cannot be satisfied. Check --gres=gpu:N against that ceiling, and check for a typo in the GRES name — the configured type is a100.

My job started, but my program says no GPU is available#

Work through these in order:

  1. Check --gres was actually in your submission script. A job submitted without --gres runs fine on a gpuqueue node with zero GPUs allocated — Slurm doesn't require a GPU request just because you picked the GPU partition.
  2. Check $CUDA_VISIBLE_DEVICES was set inside the job. Slurm sets this automatically when a GPU GRES is allocated (see Submit a GPU job). If it's empty, the allocation itself is the problem — check step 1.
  3. Check your software actually loaded a GPU-capable build. A CPU-only install of a framework that also has a GPU build will run — just without using the GPU, and usually without an error. Confirm you loaded the right module version.
  4. Check the framework can see the device from inside Python/your runtime, not just that the environment variable is set — a driver or library-version mismatch between your software build and what's installed can cause the device to be invisible to the application even though Slurm allocated it correctly.

Note Do not use --mem to fix this — a device-visibility problem is not a memory problem, and increasing system RAM does nothing for it.

My job failed with an out-of-memory error — which kind?#

A "CUDA out of memory" error and a Slurm OUT_OF_MEMORY state are not the same failure — see the comparison below.

This is the distinction that matters most, and the two are unrelated:

System memory (RAM) OOMGPU memory (VRAM) OOM
What ran outThe --mem/--mem-per-cpu you requestedThe GPU's own memory — not requested through Slurm at all
How you'll see itSlurm state OUT_OF_MEMORY in squeue/sacctAn error from your application/CUDA runtime (commonly something like "CUDA out of memory") in your job's output, not a Slurm state
What to checksacct -j <jobid> --format=State,ReqMem,MaxRSS — see Why did my job fail?Your program's own error output — Slurm has no visibility into VRAM usage
What actually fixes itIncrease --mem/--mem-per-cpu and resubmitReduce what your program asks the GPU to hold — smaller batch size, smaller model, gradient accumulation. Increasing --mem does nothing here.

Important A CUDA/GPU out-of-memory error is an application-level failure, not a Slurm resource-limit failure. Slurm does not track or enforce GPU memory on this cluster, so there is no Slurm-side setting to raise. The fix lives in your job's own parameters (batch size, model size, precision), not in your #SBATCH directives.

My job ran, but seems to be using the CPU instead of the GPU#

Usually one of the causes under "no GPU is available" above, just discovered after the fact rather than from an explicit error — many frameworks silently fall back to CPU rather than failing loudly. Check your output logs for any framework-level warning about device selection, and verify with gpustat (see Submit a GPU job) during a short test run that the GPU is actually being used.

Driver, CUDA, or library version mismatches#

If your software reports a version-incompatibility error (a CUDA runtime that doesn't match the driver, or a library built against a different CUDA version than what's loaded), that's a software environment problem, not a scheduling one. Confirm exactly which module version you loaded, and check whether a different version of the same tool is available:

bash
module avail <name>

If nothing available fits, this is a Request software situation rather than something to work around.

Getting help#

Raise it through Getting help when you've worked through the relevant section above and the job still doesn't behave as expected. Include the job ID, your submission script, and — for a GPU-visibility or version problem — the exact error text from your program's output.