Why is my job pending?
Usually nothing is wrong. Mjolnir is shared, and at any moment there are typically several hundred jobs queued behind the ones currently running. Waiting is the normal state of a healthy cluster, not a symptom.
The question worth answering is *which* kind of waiting you are looking at.
Step 1 — Get the reason#
squeue -u $USER -o "%.10i %.8T %R"The last column is the reason. In the default squeue output it is the NODELIST(REASON) column, shown in parentheses.
Step 2 — Look it up#
Wait — this is normal scheduling#
| Reason | What it means |
|---|---|
Priority | Other jobs are ahead of yours. By far the most common reason. |
Resources | Your job is next in line, but the resources it needs are not free yet. |
These need no action. Priority in particular is simply the queue doing its job: your position depends mostly on fairshare, which falls as you use the cluster and recovers as you idle. See Partitions, QoS, and fairshare.
If you want to wait less, ask for less. A job requesting fewer CPUs, less memory, or a shorter runtime is easier for the scheduler to fit into a gap and often starts sooner.
Wait — you asked for it#
| Reason | What it means |
|---|---|
Dependency | The job is waiting for another job you told it to wait for. |
JobArrayTaskLimit | Your array is running as many tasks at once as its throttle allows. |
Also normal. Dependency resolves when the job it depends on finishes — but if that job failed or was cancelled, the dependent job may wait indefinitely, so check the state of the job it is waiting for.
Change something — the request cannot be satisfied#
Reasons beginning QOSMax…, AssocMax…, or PartitionTimeLimit mean the job is asking for more than its QoS, account, or partition permits. The job will not start on its own, because nothing about waiting makes the request legal.
For example, a job asking for more CPUs than the normal QoS allows per job will be held rather than scheduled. The current limit is 48 CPUs per job and per node — Partitions, QoS, and fairshare lists all current limits.
What to do:
- Compare your
#SBATCHrequest against the current limits. - Adjust the script — fewer CPUs, less memory, a shorter
--time, or a different partition. - Cancel the held job with
scancel <jobid>and resubmit.
Important A job in this state does not fix itself. It will keep waiting until you change the request or cancel it.
Get help#
Raise it through Getting help when:
- The reason is one you cannot map to your own request
- A job with a modest request has waited far longer than comparable jobs
- The reason refers to a node or partition state rather than your job
Dependencypersists after the job it depends on has clearly finished
Include the job ID, the reason string, and your submission script.
Step 3 — Sanity-check your own request#
Before escalating, confirm the job is asking for something reasonable:
scontrol show job <jobid>Check the requested CPUs, memory, time limit, partition, and QoS against the current limits. A request that is legal but very large will simply wait longer — the scheduler has to free that much at once.
Related#
- Monitoring jobs — states and what they mean
- Submitting jobs — requesting resources sensibly
- Check job efficiency — evidence for a smaller, faster-to-schedule request
- Partitions, QoS, and fairshare — limits and priority
