Per base sequence content of Metagenomic samples
I am carrying out taxonomic classification of some metagenomic WGS samples (sequenced using Illumina NovaSeq and NexteraPE adaper libraries) using Kraken2 and I have a couple of questions related to the per base sequence content of some of these samples after quality control,
- One of the samples had this particular per base sequence content (image below). The bias observed in the content, can this be a result of duplicate reads? Note: While this particular sample had around 50% duplicate reads some samples that had similar per base sequence content but only had around 10 - 20% duplicate reads.
- Is it advisable to use such samples in taxonomic classification?
• 180 views
•
link
0 answers
No answers yet.
Log in to answer this question.
It is possible (assuming nothing else has gone wrong in the experiment) that your samples have low diversity. That combined with ample amplification will result in high duplicates.