Submitting sequencing data to ENA and NCBI SRA

If your sequencing data is already on Mjolnir, you can upload it to a public archive from here. The submission clients are installed as modules, and Mjolnir can reach both archives' upload services.

This page covers the Mjolnir side: which client to load, what to run, and where to run it. What each archive requires of your submission — study and sample metadata, manifests, checklists, accessions — is the archive's own documentation, and this page links to it rather than repeating it.

Note For transferring data to or from ERDA, consult the official UCPH ERDA User Guide.

Pick the archive first — the two workflows differ#

ENA and NCBI SRA use different submission workflows.

AspectEuropean Nucleotide Archive (ENA)NCBI Sequence Read Archive (SRA)
Where metadata is enteredWebin service, or generated by Webin-CLIWeb Submission Portal wizard
Order of workUpload and submit can be one stepMetadata first, files afterwards
Recommended Mjolnir clientena-webin-cliaspera-cli (ascp) or lftp
Upload and submissionWebin-CLI does bothAlways separate steps

If you only need one archive, read only its section below.

Before you start#

You need an account with the archive. Registration is the archive's process, not Mjolnir's:

ENA requires an MD5 checksum for every file and re-computes it on arrival to confirm the transfer did not alter the file. Generate them on Mjolnir before you upload:

bash
md5sum *.fastq.gz > md5sums.txt

That writes one line per file — the checksum, then the filename — into md5sums.txt. You can re-check your local copies at any time with:

bash
md5sum -c md5sums.txt

Each file is reported as OK or FAILED.

Note NCBI's file-upload documentation does not state the same explicit MD5 requirement as ENA. Generating checksums is still a sensible habit for your own verification, but follow NCBI's instructions for what its submission actually requires.

If you need to compress data first, that is CPU-heavy work — send it to a compute job rather than running it on the login node. See Data transfer for the pattern.

ENA: the integrated route (Webin-CLI)#

Webin-CLI validates your files, uploads them, and submits them — you do not pre-upload separately. It provides the most integrated ENA route available on Mjolnir.

bash
module load ena-webin-cli/9.0.3

The module brings its own Java runtime, so there is nothing else to load or configure.

Check what it expects:

bash
ena-webin-cli -help

Webin-CLI works in contexts — reads, genome, transcriptome, sequence, polysample, taxrefset — and each context takes a manifest file listing your data files and the metadata fields for that submission type. The fields differ per context, so build your manifest from ENA's own reference rather than from an example here:

The usual sequence is to validate first, then submit the same manifest:

bash
ena-webin-cli -context=reads -manifest=manifest.txt -userName=Webin-XXXXX -passwordFile=~/.webin-pass -validate
ena-webin-cli -context=reads -manifest=manifest.txt -userName=Webin-XXXXX -passwordFile=~/.webin-pass -submit

Validation reports are written under the output directory, so you can read what failed before anything is sent.

Tip Add -test to run against ENA's test service instead of the live archive. It is the safe way to rehearse a first submission.

ENA: staging files only (lftp)#

If you would rather upload files now and complete the metadata submission separately through Webin, upload them into your private Webin file upload area over FTP.

bash
module load lftp/4.9.3
lftp webin2.ebi.ac.uk -u Webin-XXXXX

lftp prompts for your Webin password, then gives you an interactive session — put, mput, ls, bye. It can resume an interrupted upload rather than starting over.

Important This only stages the files. It does not submit them. Your submission is not made until you complete the metadata step through the Webin service — see Uploading files to ENA.

ENA expects staged files to be submitted reasonably promptly: its documentation states that a file is not expected to sit in the upload area for more than about two months before you instruct it to be submitted. Stage when you are ready to finish the job, not months ahead.

NCBI SRA#

SRA works the other way round: the submission is created in the browser first, and the file upload is a later step that points at files you have already transferred.

  1. In the Submission Portal, register or select your BioProject and BioSamples and complete the SRA metadata.
  2. Request your personal account folder (a "preload" folder). NCBI recommends this because it lets you upload files independently of the wizard.
  3. Upload your files from Mjolnir into that folder, using either method below.
  4. Return to the Submission Portal and select the folder to finish the submission.

Your account folder name and your upload credentials both come from the Submission Portal. Use the values it shows you.

Upload with Aspera#

bash
module load aspera-cli/3.9.6

NCBI documents the transfer as:

bash
ascp -i <path/to/key-file> -QT -l 100m -k1 -d <folder-with-files> subasp@upload.ncbi.nlm.nih.gov:uploads/<your-account-folder>

The private key file is issued to you through the Submission Portal. -k1 is NCBI's documented way to enable resuming a partial transfer, -l 100m caps the rate, and -d creates the target directory. The current flags and destination are in NCBI's file-upload documentation.

Upload with FTP#

bash
module load lftp/4.9.3
lftp ftp-private.ncbi.nlm.nih.gov -u <username-from-submission-portal>

Then change into your account folder, make a subfolder for this submission, and put the files — the exact folder path is the one the Submission Portal gives you.

Important sra-tools (prefetch, fasterq-dump) is for retrieving data from SRA. It is not a submission tool and cannot upload. Use Aspera or FTP as above.

If an Aspera transfer connects but does not progress#

Aspera authenticates over TCP and then moves the actual data over UDP port 33001. Mjolnir's outbound path for that UDP data channel has not yet been confirmed by a completed authenticated transfer, so if ascp authenticates and then sits at zero bytes with no progress, do not spend long on it:

  • Use the archive's FTP route instead — outbound FTP from Mjolnir, including the passive data connection, has been verified working.
  • Tell us through the Support Center so we can confirm the path and update this page.

Ordinary FTP uploads are unaffected by this.

Large or long-running uploads#

An ordinary upload runs from the login node, and that is fine — moving files is what the login node is for. You do not need a Slurm job to submit a few runs to an archive.

For a large, long-running upload — moving a substantial dataset rather than an ordinary day-to-day transfer — use Mjolnir's dedicated filetransfer partition and QoS, the same facility described in Data transfer. It allows up to 4 days of runtime, and the job keeps running after you disconnect.

The submission clients and the archive endpoints are both available from those nodes:

bash
#!/bin/bash
#SBATCH --job-name=ena-upload
#SBATCH --account=mjolnir
#SBATCH --partition=filetransfer
#SBATCH --qos=filetransfer
#SBATCH --cpus-per-task=1
#SBATCH --mem=20G
#SBATCH --time=1-00:00:00
#SBATCH --output=ena-upload-%j.out

module load ena-webin-cli/9.0.3
ena-webin-cli -context=reads -manifest=manifest.txt -userName=Webin-XXXXX -passwordFile=$HOME/.webin-pass -submit

Save the script as ena-upload.sh and submit it from your Mjolnir login session:

bash
sbatch ena-upload.sh

Then watch it like any other job:

bash
squeue -u $USER --partition=filetransfer

Important The filetransfer QoS is limited to 1 CPU and 20 GB per job, and is for transfer work only. If your job stays pending, see Why is my job pending? — asking for more than those limits is a common cause.

Keeping your credentials safe#

Archive credentials are personal, and Mjolnir project directories are shared.

  • Never put an archive password or a private key in a project, scratch, or otherwise group-readable directory.
  • Never store passwords or private keys in job scripts. Job scripts may be copied, shared, or committed.
  • Use the mechanisms the clients provide. Webin-CLI accepts -passwordFile=<file> or -passwordEnv=<VARIABLE> instead of -password on the command line, which also keeps the password out of your shell history and out of squeue output.
  • Keep any password file and the NCBI Aspera key private to your own account:
bash
chmod 600 ~/.webin-pass
chmod 600 ~/aspera-key-file
  • If a credential is ever exposed, change it with the archive rather than hoping it went unnoticed.