This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Very high number of unique sequences in PacBio full-length 16S CCS data using DADA2

I am analysing PacBio Sequel II full-length 16S rRNA CCS reads (~1450 bp) using the DADA2 long-read workflow and observing an unusually high number of unique sequences.

Almost all reads appear unique (e.g., ~11,200 unique reads from ~12,300 total reads). After denoising (learnErrors, dada, chimera removal), only a small number of reads remain.

Is such a high unique/read ratio normal for PacBio full-length 16S CCS data? Could this be related to sequence orientation, primer trimming, or filtering parameters?

Any suggestions for diagnosing or resolving this issue would be appreciated.

pacbio dada2 microbiome amplicon 16s

0 answers

No answers yet.

Log in to answer this question.