How Mjolnir works
Mjolnir is a shared system, so using it works differently from using your own computer: you log in to one machine, but your analyses run on others. This page is the single picture of how those pieces fit together. It is conceptual — every page it links to has the exact detail.
The path from login to computation#
- From your own computer, connect to the UCPH VPN, then SSH to Mjolnir.
- You land on the gate node. This is where you prepare work.
- You submit a job describing what to run and what it needs.
- Slurm decides when it runs and which machine it runs on.
- Your job executes on a compute node, reading and writing the same shared storage you can see from the gate node.
The step people are most often surprised by is 3. You do not start your analysis yourself — you describe it and hand it to the scheduler.
Gate node#
The gate node is the entrance to Mjolnir and the only machine you log in to directly. It is shared by everyone connected at that moment.
It is for preparing and managing work:
- Logging in, and moving around the filesystem
- Editing scripts and setting up environments
- Copying files in and out
- Submitting jobs, and checking on them afterwards
It is not for running analyses. A heavy process there slows down every other person's session, including people only trying to check on a job. Running computation on the gate node is not permitted.
Slurm#
Because many people share a limited number of machines, something has to decide who gets what, and when. That is Slurm, the scheduler.
You submit a job — a short script saying what to run, how many CPUs it needs, how much memory, and for how long. Slurm queues it, and starts it on a suitable machine as soon as one is free and it is your turn. Asking for roughly what you actually need is what makes this work well for you and for everyone else.
Compute nodes and partitions#
Compute nodes are where work actually runs. They are the large machines — many CPU cores, a lot of memory, and GPUs on some of them.
Mjolnir's compute nodes are one shared fleet. Partitions are not separate clusters; they are different ways of asking for the same machines:
cpuqueue— general work, and the defaultgpuqueue— the machines that have GPUslazyqueue— the same machines again, at low priority, for work that can be interrupted and restarted
This is worth internalising early, because it explains why a job can wait even though "another partition looks free": much of the time, it is the same hardware.
Shared storage#
Your files live on shared storage — chiefly /home for personal configuration, /projects for the real work of your group, and /datasets for read-only reference databases.
The important property is this: the gate node and every compute node see the same files, at the same paths. A file you create in your project folder on the gate node is already visible to your job, under exactly the same path.
So you normally do not copy data onto individual compute nodes before running, and you do not fetch results back off one afterwards — it is the same filesystem either way. Staging to node-local temporary storage is occasionally worth it for specific I/O-heavy patterns, but it is an optimisation to reach for deliberately, not the normal way to work.
Where to go next#
| To do this | Read |
|---|---|
| Connect for the first time | VPN and SSH access |
| Walk through a first job, start to finish | First 15 minutes on Mjolnir |
| Understand what belongs where | Login nodes vs compute nodes |
| Write and submit a job script | Submitting jobs |
| See the exact partitions, limits, and priority rules | Partitions, QoS, and fairshare |
| Decide where your files belong | Where should my files go? |
