Seconding this. Please don't overinterpret these fastqc results. It's really just a crude QC check. Go along with your analysis.
Hi everyone, I was recently performing a fastQC after adapter timing with TrimGalore, and found a strange overrepresented sequence in read pair 2 for most of my samples: 'GTAAAAGGTAGCAATAGCTTTAAGCCAAGAAATTGTTCTCAGAAATGGCT'
Has anyone come across it before and knows what it relates to?
Background info: I used the following parameters for trim_galore: ' --cores 4 --trim-n --length 36 --paired --retain_unpaired '
Thank you!
1 answer
It is 0.1% of the total sequence. Nothing you likely need to worry about at this point.
I did see that one. I do wonder why these seqs would be over-represented in R2 and not R1... anyhow, as all of you mention, I will go ahead, now that I know where they come from. I thought they were some sort of weird adapter derived sequence but they actually come from the chloroplast...
Log in to answer this question.
Hi, could you please provide further information on the type of data your reads originate from? Is it RNA-Seq or DNA-Seq - which species/genus does it originate from?
I just blasted it and found a Nicotiana attenuata cloroplast-based predicted protein. Would that fir your data?
It would, thanks for your input!