Hello everyone,
I am new to sequencing and DNA analysis, as well as bioinformatics in general. I am working with a very small insect. I extracted DNA using a Qiagen kit and sent the samples for Sanger sequencing. I am only targeting a specific region, the mitochondrial COI gene, which has been amplified and sequenced. The purpose is simply to identify insect species using DNA sequence analysis through BLAST searches on the NCBI and BOLD databases. Now I have received the sequencing results, but I don’t know how to properly analyze the data.
I am familiar with CLC software and have tried using it, but during training or repeated analyses, the settings seem to change and I am not getting consistent or correct results. When I BLAST the sequences on NCBI, the query coverage and identity percentages are always very low, so I am not confident about the results.
I also tried analyzing the data using the Galaxy server, but I find it very complicated as a beginner.
Could someone please guide me on the correct workflow for analyzing sequencing data from insect DNA, or suggest beginner-friendly tools and steps? Any help would be greatly appreciated. Thank you.
2 answers
Low coverage and low identity together usually means the read is bad, not that the analysis is wrong. A clean COI barcode hits GenBank or BOLD at 95-100% over most of its length even when your exact species isn't in there, because congeners still match. So I'd stop adjusting software settings and look at the trace itself.
Two things to check on the chromatogram. The first 20-40 bases and the tail are always unreliable, so if you're BLASTing the raw read with primer sequence still attached, coverage and identity both drop for reasons that have nothing to do with your insect - trim to the clean window and BLAST that. And doubled peaks running through the whole trace mean you amplified more than one template, which happens a lot with whole-body extractions of small insects. That one is a bench problem and no amount of analysis will fix it.
Disclosure, I work on SeqBench: https://seqbench.com/tools/sanger-trace-viewer opens an .ab1 in the browser, shows per-base Phred quality and exports the trimmed read as FASTA. Nothing to install or misconfigure, which sounds like the CLC issue you're hitting. Any ab1 viewer does the same job.
I think the main issue pertaining to CLC is this "I am familiar with CLC software and have tried using it, but during training or repeated analyses, the settings seem to change and I am not getting consistent or correct results. When I BLAST the sequences on NCBI, the query coverage and identity percentages are always very low, so I am not confident about the results. "
The above user brings up some good points:
"Low coverage and low identity together usually means the read is bad, not that the analysis is wrong. A clean COI barcode hits GenBank or BOLD at 95-100% over most of its length even when your exact species isn't in there, because congeners still match. So I'd stop adjusting software settings and look at the trace itself. Two things to check on the chromatogram. The first 20-40 bases and the tail are always unreliable, so if you're BLASTing the raw read with primer sequence still attached, coverage and identity both drop for reasons that have nothing to do with your insect - trim to the clean window and BLAST that. And doubled peaks running through the whole trace mean you amplified more than one template, which happens a lot with whole-body extractions of small insects. That one is a bench problem and no amount of analysis will fix it."
Were the amplicons cloned into a vector before sequencing? This can help ensure that a single template is sequenced.
Read QC will be important. If raw reads are being used as input to BLAST, then the vector is still on the reads and it will need to be trimmed away. This can be done in CLC using the trim sequences tool: https://resources.qiagenbioinformatics.com/manuals/clcgenomicsworkbench/current/index.php?manual=Trim_sequences.html
The "Trim contamination from sequences " option can be used to specify the vector sequences used for cloning. Trim using quality scores will be useful in removing the low quality base calls near the ends of the reads (which is typical for sanger reads). CLC uses the same BLAST algorithm as NCBI. There are options to change the alignment parameters as needed:
It might be worth running the Trim and Map Sanger sequences template workflow:
This workflow will remove low quality base calls/vector from the reads and map them to a reference sequence. You can use the mitochondrial COI gene as the reference. Once the alignment finishes, you can inspect the read mapping file to see if your primers are amplifying the right product.
I hope this helps, please reach out to TS-Bioinformatics@qiagen.com with any questions
Log in to answer this question.
What are you specifically trying to do with the Sanger sequencing?
I want identify insect species using DNA sequence analysis through BLAST searches on NCBI and the BOLD database.
It is unlikely that you are trying to sequence the genome of insect of interest using
sangersequencing. Are you looking at region(s) of interest (genes) that have been amplified and then sequenced? You will need to provide additional information about the experiment to get useful answers.You can use these two programs mentioned in answers here as user friendly local options to analyze sanger data : Sanger sequencing data analysis
No, I am not trying to sequence the whole genome. I am only targeting a specific region, the mitochondrial COI gene, which has been amplified and sequenced. The purpose is simply to identify insect species using DNA sequence analysis through BLAST searches on the NCBI and BOLD databases.
It is possible that something failed during the amplification. You will need to check that.
Are you using a single set of primers? If so you should be easily able to do a multiple sequence alignment of your sequences and generate a consensus. You can then use that to BLAST at NCBI limiting the search to e.g.
insects (taxid:6960).