Why did my job fail?

Your job failed, or finished without producing what you expected. Work through the evidence in order — the state usually tells you which of a few very different problems you have.

The four outcomes you are most likely to be looking at are a failed job, an out-of-memory (OOM) kill, a timeout, or a cancelled job. They need completely different responses.

Step 1 — Ask Slurm what happened#

bash
sacct -j <jobid> --format=JobID,JobName,State,ExitCode,Elapsed,ReqMem,MaxRSS,AllocCPUS,NodeList

The State column is the first thing to read. Everything else interprets it.

Step 2 — Read the state#

StateWhat it means
COMPLETEDSlurm ran the script to the end successfully. If the result is wrong, the problem is in your commands, not the scheduler.
FAILEDYour script exited with a non-zero status — usually an error inside the program.
OUT_OF_MEMORYThe job exceeded the memory it requested and was killed.
TIMEOUTThe job hit its --time limit and was killed.
CANCELLEDSomeone or something cancelled it — often the user, sometimes an administrator.

Note FAILED is by far the most common non-successful outcome, and it almost always means the program inside your job returned an error — not that the cluster malfunctioned. Read your error file before assuming otherwise.

Step 3 — Read your own output#

Slurm records the outcome; your program explains it. Look at the files named by --output and --error in your script:

bash
cat myjob-1234567.err
tail -50 myjob-1234567.out

If you did not set --error, errors are mixed into the output file.

An empty output file usually means the job failed before your commands ran — a bad path, a missing module, or a script that could not be read.

Common causes, by state#

FAILED#

Most often one of:

  • command not found — the module was not loaded *inside the job*. Loading it in your terminal does not carry over. See Find and load software.
  • No such file or directory — a relative path that was valid where you submitted but not where the job ran. Use absolute paths, or cd "$SLURM_SUBMIT_DIR" first.
  • Permission denied — the job cannot read or write where you told it to. Check the location is one you have access to (Where should my files go?).
  • A genuine error in the program — bad arguments, malformed input, unmet assumptions.

What to change: fix the cause shown in the error file and resubmit. Retrying unchanged will fail identically.

OUT_OF_MEMORY (an OOM kill)#

The job asked for a certain amount of memory and tried to use more, so it was killed. This is what people mean by "the job got OOM-killed".

Compare what you asked for against what it actually reached:

bash
sacct -j <jobid> --format=JobID,State,ReqMem,MaxRSS,AllocCPUS

ReqMem is your request; MaxRSS is the peak memory recorded. If MaxRSS is at or near your request, memory is the problem.

What to change: raise the request and resubmit — for example --mem-per-cpu=4G instead of 2G, or set a total with --mem=32G. Increase deliberately rather than requesting an enormous amount: over-requesting makes jobs wait longer and counts against your fairshare (Submitting jobs). See how much memory you can actually request if you are raising it a lot.

If a job needs far more memory than expected, that is worth understanding rather than papering over — an accidental full-file read or an unbounded data structure is a common cause.

Note A 2025 Mjolnir announcement introduced PSS (Proportional Set Size) memory reporting, which attributes shared memory more fairly between processes, so figures may read lower than older accounting for the same work. Either way, treat MaxRSS as a practical guide for sizing your next request rather than an exact measurement.

TIMEOUT#

The job reached its --time limit and was killed. Anything not yet written is lost.

What to check: how long it actually ran (Elapsed) against what you asked for. If they match, the limit was simply too short.

What to change: raise --time to a realistic estimate of the real runtime, plus some margin. Do not simply request the partition maximum — a shorter, honest limit is easier for the scheduler to fit into a gap and often starts sooner. Partition maximums are listed in Partitions, QoS, and fairshare; a request above the maximum will not start at all.

If the work genuinely does not fit, consider splitting it, checkpointing it, or running it as a job array.

CANCELLED#

Usually somebody ran scancel — often you. It can also be an administrator, for example during maintenance. If you did not cancel it and the reason is not obvious, ask.

Step 4 — Understand the exit code#

sacct shows ExitCode as two numbers separated by a colon, such as 0:0 or 1:0.

  • The first number is the program's exit status. 0 means success; anything else is the error code the program itself returned.
  • The second number is the signal that killed the job, if any. A non-zero value here means the job was terminated rather than exiting on its own — which is what you see for memory and time-limit kills.

For practical purposes: a non-zero first number sends you to your error file; a non-zero second number sends you to the state, which will usually be OUT_OF_MEMORY or TIMEOUT.

When to ask for help#

Raise it through Getting help when:

  • The state and the evidence disagree with each other
  • The same job succeeds sometimes and fails other times with no change
  • The error points at the system rather than your work
  • You cannot map the failure to anything in your script

Include the job ID, the state and exit code, the relevant lines from your error file, and your submission script. That is usually enough to answer in one exchange.