This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Per base sequence content of Metagenomic samples

I am carrying out taxonomic classification of some metagenomic WGS samples (sequenced using Illumina NovaSeq and NexteraPE adaper libraries) using Kraken2 and I have a couple of questions related to the per base sequence content of some of these samples after quality control,

  1. One of the samples had this particular per base sequence content (image below). The bias observed in the content, can this be a result of duplicate reads? Note: While this particular sample had around 50% duplicate reads some samples that had similar per base sequence content but only had around 10 - 20% duplicate reads.

Per base sequence content of metagenomic sample

  1. Is it advisable to use such samples in taxonomic classification?
metagenomics ngs fastqc

duplicate reads.

It is possible (assuming nothing else has gone wrong in the experiment) that your samples have low diversity. That combined with ample amplification will result in high duplicates.

0 answers

No answers yet.

Log in to answer this question.