This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Prefetch error while downloading data from sequence read archive

Hi! I'm using prefetch command from SRA Toolkit on HPC cluster to download SRA data, but got stuck into error:

2020-03-26T16:34:59 prefetch.2.10.0 int: transfer incomplete while reading file within network system module - Cannot KStreamRead: https://sra-downloadb.be-md.ncbi.nlm.nih.gov/sos2/sra-pub-run-13/SRRXXXXXXX/SRRXXXXXX.X

How can i fix that?

software error

Thaks for your reply. I don't run them in parallel, but sequentially one by one inside one for loop. Before prefetch running i've executed:

/home/sratoolkit.2.10.0-ubuntu64/bin/vdb-config -s /http/timeout/read=1000000000

It seems that above command helps but not in all cases.

Besides, i should to note that initially i have approximately 60 accessions and got error with prefetch only for 11 of them (everything were fine with the rest ones). Then i ran again exactly the same prefetch command for these 11 accessions and downloaded half of them (for 6 got error).

Alternatively, you can also search for your accessions via https://sra-explorer.info/ and it will produce several download links you can try including direct download from the ftp servers from NCBI.

This tutorial here covers download directly in fastq format from ENA but be aware that ENA is performing maintenance work right now so download might not be possible. Fast download of FASTQ files from the European Nucleotide Archive (ENA)

Thanks a lot. Will try different solutions.

0 answers

No answers yet.

Log in to answer this question.