This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Strange overrepresented sequence

Hi everyone, I was recently performing a fastQC after adapter timing with TrimGalore, and found a strange overrepresented sequence in read pair 2 for most of my samples: 'GTAAAAGGTAGCAATAGCTTTAAGCCAAGAAATTGTTCTCAGAAATGGCT'

Overrepresented sequence in fastQC analysis

Has anyone come across it before and knows what it relates to?

Background info: I used the following parameters for trim_galore: ' --cores 4 --trim-n --length 36 --paired --retain_unpaired '

Thank you!

fastqc adapter rnaseq

Hi, could you please provide further information on the type of data your reads originate from? Is it RNA-Seq or DNA-Seq - which species/genus does it originate from?

I just blasted it and found a Nicotiana attenuata cloroplast-based predicted protein. Would that fir your data?

It would, thanks for your input!

1 answer

It is 0.1% of the total sequence. Nothing you likely need to worry about at this point.

Seconding this. Please don't overinterpret these fastqc results. It's really just a crude QC check. Go along with your analysis.

I did see that one. I do wonder why these seqs would be over-represented in R2 and not R1... anyhow, as all of you mention, I will go ahead, now that I know where they come from. I thought they were some sort of weird adapter derived sequence but they actually come from the chloroplast...

Log in to answer this question.