Data transfer
Ordinary transfers: workstation ↔ Mjolnir#
For everyday transfers — a few files, a results folder, something you're actively working on — connect over SSH from your own computer. This does not need a Slurm job.
# from your computer, copy a local file to your project folder
scp local-file.txt <KU-ID>@mjolnirgate.unicph.domain:/projects/<project>/people/<KU-ID>/
# from your computer, copy a whole directory back down
scp -r <KU-ID>@mjolnirgate.unicph.domain:/projects/<project>/people/<KU-ID>/results ./rsync is the better default once a transfer might get interrupted or you need to run it again — unlike scp, it can resume and will skip files that already copied successfully:
rsync -avh --progress local-folder/ <KU-ID>@mjolnirgate.unicph.domain:/projects/<project>/people/<KU-ID>/local-folder/-apreserves permissions, timestamps, and symlinks-vshows what's being copied-hprints sizes in human-readable form--progressshows progress as it runs
Warning Note the trailing slash on the source (local-folder/). With it, rsync copies the *contents* of local-folder into the destination; without it, rsync creates local-folder itself inside the destination. Get this backwards and you can end up with files nested one level deeper than you expected.
If you prefer a graphical client, FileZilla (cross-platform) and WinSCP or PuTTY-based tools connect the same way — host mjolnirgate.unicph.domain, your KU-ID, your KU password, port 22.
This is still login-node activity#
An ordinary transfer runs through the login node, and that's fine — moving files is exactly what the login node is for (see Login nodes vs compute nodes). What doesn't belong there is the CPU-heavy work that sometimes rides along with a transfer, like compressing a large dataset first.
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=2G
#SBATCH --time=02:00:00
pigz -p $SLURM_CPUS_PER_TASK -9 large_file.txtpigz is a parallel gzip — submit it as a job rather than running it on the login node.
Large migrations: the filetransfer QoS#
For large, long-running transfers — moving a substantial dataset between locations, not day-to-day file copying — Mjolnir provides a dedicated filetransfer partition and QoS built for exactly that, verified live:
| Setting | Value |
|---|---|
| Partition | filetransfer |
| QoS | filetransfer |
| Max runtime | 4 days |
| CPU / memory per job | 1 CPU, 20 GB |
| Requires | the mjolnir account and the filetransfer QoS |
It exists so a large transfer gets priority scheduling without competing with — or being limited by — ordinary compute jobs.
sbatch --partition=filetransfer \
--qos=filetransfer \
--cpus-per-task=1 \
--mem=20G \
--time=1-00:00:00 \
--wrap="rsync -avh /path/to/source/ /path/to/destination/"Important This QoS is for transfer commands only — rsync, scp, sftp and similar. It is not a general-purpose compute QoS; running compute work through it is a misuse of a shared resource intended for everyone's large transfers.
Check on it the same way as any job:
squeue -u $USER --partition=filetransferIf it stays pending, see Why is my job pending? — a common cause here is requesting more than the 1 CPU / 20 GB the QoS allows.
Choosing between the two#
| Situation | Use |
|---|---|
| A few files, actively working | Ordinary scp/rsync over SSH |
| A large one-off migration | filetransfer QoS job |
| Compressing data before either | A cpuqueue job — never the login node |
Related#
- Submitting sequencing data to ENA and NCBI SRA — uploading from Mjolnir to a public archive
- Where should my files go? — destinations for what you're moving
- Login nodes vs compute nodes
- Why is my job pending?
- Getting help
