Why is my job pending?

Usually nothing is wrong. Mjolnir is shared, and at any moment there are typically several hundred jobs queued behind the ones currently running. Waiting is the normal state of a healthy cluster, not a symptom.

The question worth answering is *which* kind of waiting you are looking at.

Step 1 — Get the reason#

bash
squeue -u $USER -o "%.10i %.8T %R"

The last column is the reason. In the default squeue output it is the NODELIST(REASON) column, shown in parentheses.

Step 2 — Look it up#

Wait — this is normal scheduling#

ReasonWhat it means
PriorityOther jobs are ahead of yours. By far the most common reason.
ResourcesYour job is next in line, but the resources it needs are not free yet.

These need no action. Priority in particular is simply the queue doing its job: your position depends mostly on fairshare, which falls as you use the cluster and recovers as you idle. See Partitions, QoS, and fairshare.

If you want to wait less, ask for less. A job requesting fewer CPUs, less memory, or a shorter runtime is easier for the scheduler to fit into a gap and often starts sooner.

Wait — you asked for it#

ReasonWhat it means
DependencyThe job is waiting for another job you told it to wait for.
JobArrayTaskLimitYour array is running as many tasks at once as its throttle allows.

Also normal. Dependency resolves when the job it depends on finishes — but if that job failed or was cancelled, the dependent job may wait indefinitely, so check the state of the job it is waiting for.

Change something — the request cannot be satisfied#

Reasons beginning QOSMax…, AssocMax…, or PartitionTimeLimit mean the job is asking for more than its QoS, account, or partition permits. The job will not start on its own, because nothing about waiting makes the request legal.

For example, a job asking for more CPUs than the normal QoS allows per job will be held rather than scheduled. The current limit is 48 CPUs per job and per node — Partitions, QoS, and fairshare lists all current limits.

What to do:

  1. Compare your #SBATCH request against the current limits.
  2. Adjust the script — fewer CPUs, less memory, a shorter --time, or a different partition.
  3. Cancel the held job with scancel <jobid> and resubmit.

Important A job in this state does not fix itself. It will keep waiting until you change the request or cancel it.

Get help#

Raise it through Getting help when:

  • The reason is one you cannot map to your own request
  • A job with a modest request has waited far longer than comparable jobs
  • The reason refers to a node or partition state rather than your job
  • Dependency persists after the job it depends on has clearly finished

Include the job ID, the reason string, and your submission script.

Step 3 — Sanity-check your own request#

Before escalating, confirm the job is asking for something reasonable:

bash
scontrol show job <jobid>

Check the requested CPUs, memory, time limit, partition, and QoS against the current limits. A request that is legal but very large will simply wait longer — the scheduler has to free that much at once.