I run fastqc on Illumina fastq files from miseq and found a very high level of sequence duplicate as reported elsewhere. Here, I just want to understand how the total duplicate percentage is calculated. The output from fastqc_data.txt is:
Sequence Duplication Levels fail
Total Duplicate Percentage 90.87717882197605 Duplication Level Relative count
1 100.0
2 17.39881642975786
3 8.231572445683474
4 4.841954142913186
5 2.9630989841995876
6 1.9648267630438956
7 1.3296385019239276
8 0.9081272379744089
9 0.6627325615364712
10++ 3.6083033545619205
Where is 90.87717882197605 coming from? Thank in advance.
1 answer
If I recall this correctly those percentages are computed relative to the number of unique reads. See how the first number is 100%. The total duplicate percentage is relative to the total number of reads in the sample.
What this means is that even though the number of reads that are duplicated more than 10 times is only 3% some of these are duplicated at very high rates, tens of thousands of times and thus produce more than 90% of the data.
Thank you all for the prompt responses.
Log in to answer this question.