Hi Thomas, thanks for the quick response.
Reads have between 70-150bp of length
I was using the following command: $bwa mem -t $threads $ref_fasta ${filename}_1.fastq.gz ${filename}_2.fastq.gz > ${filename}.sam
Hi everyone.
I have paired end sequencing data (Illumina) and there are specific regions in the genome I am interested in. I aligned the samples using bwa to the hg38 fasta and it took 19h to align and generate the SAM file. I wanted to speed up this process so I filtered the hg38 fasta file to only contain positions 5000 base pairs to the left or right of my regions of interested. When aligning to the new hg38_small.fasta it took almost 30h.
I was wondering if anyone has any tip or knows about a better approach for doing this?
Thanks, Francisco
Generally aligning to heavily masked genomes like this is not a good idea as you are at risk of generating spurious alignments
How long are your reads? What exact bwa commands are you using? Depending on your read length, it may be a better idea to use bwa-mem
Using the wrong bwa tools for your read length will produce long run times
Hi Thomas, thanks for the quick response.
Reads have between 70-150bp of length
I was using the following command: $bwa mem -t $threads $ref_fasta ${filename}_1.fastq.gz ${filename}_2.fastq.gz > ${filename}.sam
Hmm, I am not sure what the issue is here exactly, as a standard alignment process like this shouldn't take 19 hours
What are the size of your fastq files?
Log in to answer this question.
please, don't. Exome Sequencing: Masking The Non-Genic Sequences ?