Hey Antonio, thank you so much for your reply, I appreciate your effort!
The data was on array express already split into forward and reverse files:http://www.ebi.ac.uk/ena/data/view/SRS1019180
I only did fastqc on each file, the forward and reverse and got good quality reports:
Only per base sequence content, per sequence GC content, sequence duplication levels and Kmers were marked with an x. The rest were marked as correct.
So I did tophat right away without any filtering and got the aformentioned alignment summary.
What do you suggest I should do?
Thank you,
Sarah
What's the length distribution of left and right reads according to fastqc report?
Hey Noolean, thank you so much for your reply, I appreciate your effort!
The data was on array express already split into forward and reverse files:http://www.ebi.ac.uk/ena/data/view/SRS1019180
I only did fastqc on each file, the forward and reverse and got good quality reports:
Only per base sequence content, per sequence GC content, sequence duplication levels and Kmers were marked with an x. The rest were marked as correct.
The ength distribution of left and right reads according to fastqc report is exactly the same:
from 98 to 100 with a peak at 99.
What do you suggest I should do?
Thank you so much,
Sarah
Check error rates, R2 reads are usually lower in quality than R1, trimming for quality might help.
Hey apelin20, thank you so much for your reply, I appreciate your effort!
The data was on array express already split into forward and reverse files:http://www.ebi.ac.uk/ena/data/view/SRS1019180
I only did fastqc on each file, the forward and reverse and got good quality reports:
Only per base sequence content, per sequence GC content, sequence duplication levels and Kmers were marked with an x. The rest were marked as correct.
What do you mean by error rates? Do you think I need trimming? There are no signs of adapters or primers in the overrepresented sequences.
Thank you,
Sarah