Data transfer

Ordinary transfers: workstation ↔ Mjolnir#

For everyday transfers — a few files, a results folder, something you're actively working on — connect over SSH from your own computer. This does not need a Slurm job.

bash
# from your computer, copy a local file to your project folder
scp local-file.txt <KU-ID>@mjolnirgate.unicph.domain:/projects/<project>/people/<KU-ID>/

# from your computer, copy a whole directory back down
scp -r <KU-ID>@mjolnirgate.unicph.domain:/projects/<project>/people/<KU-ID>/results ./

rsync is the better default once a transfer might get interrupted or you need to run it again — unlike scp, it can resume and will skip files that already copied successfully:

bash
rsync -avh --progress local-folder/ <KU-ID>@mjolnirgate.unicph.domain:/projects/<project>/people/<KU-ID>/local-folder/
  • -a preserves permissions, timestamps, and symlinks
  • -v shows what's being copied
  • -h prints sizes in human-readable form
  • --progress shows progress as it runs

Warning Note the trailing slash on the source (local-folder/). With it, rsync copies the *contents* of local-folder into the destination; without it, rsync creates local-folder itself inside the destination. Get this backwards and you can end up with files nested one level deeper than you expected.

If you prefer a graphical client, FileZilla (cross-platform) and WinSCP or PuTTY-based tools connect the same way — host mjolnirgate.unicph.domain, your KU-ID, your KU password, port 22.

This is still login-node activity#

An ordinary transfer runs through the login node, and that's fine — moving files is exactly what the login node is for (see Login nodes vs compute nodes). What doesn't belong there is the CPU-heavy work that sometimes rides along with a transfer, like compressing a large dataset first.

bash
#SBATCH --partition=cpuqueue
#SBATCH --qos=normal
#SBATCH --nodes=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=2G
#SBATCH --time=02:00:00

pigz -p $SLURM_CPUS_PER_TASK -9 large_file.txt

pigz is a parallel gzip — submit it as a job rather than running it on the login node.

Large migrations: the filetransfer QoS#

For large, long-running transfers — moving a substantial dataset between locations, not day-to-day file copying — Mjolnir provides a dedicated filetransfer partition and QoS built for exactly that, verified live:

SettingValue
Partitionfiletransfer
QoSfiletransfer
Max runtime4 days
CPU / memory per job1 CPU, 20 GB
Requiresthe mjolnir account and the filetransfer QoS

It exists so a large transfer gets priority scheduling without competing with — or being limited by — ordinary compute jobs.

bash
sbatch --partition=filetransfer \
       --qos=filetransfer \
       --cpus-per-task=1 \
       --mem=20G \
       --time=1-00:00:00 \
       --wrap="rsync -avh /path/to/source/ /path/to/destination/"

Important This QoS is for transfer commands only — rsync, scp, sftp and similar. It is not a general-purpose compute QoS; running compute work through it is a misuse of a shared resource intended for everyone's large transfers.

Check on it the same way as any job:

bash
squeue -u $USER --partition=filetransfer

If it stays pending, see Why is my job pending? — a common cause here is requesting more than the 1 CPU / 20 GB the QoS allows.

Choosing between the two#

SituationUse
A few files, actively workingOrdinary scp/rsync over SSH
A large one-off migrationfiletransfer QoS job
Compressing data before eitherA cpuqueue job — never the login node