Clustering ilumina reads with different lengths
Hi,
I have a set of fastq reads that I would like to cluster, independent of read length.
Having the initial data:
AAAAAAAAAAAAAAAAAAAAAAAAA
AAAAAAAAAAAAAAA
AAAAAAA
BBBBBBBBBBBBBBBBBBBBBBBBB
BBBBBBBBBBBBB
BBBBBB
I would like the output of my data to be:
AAAAAAAAAAAAAAAAAAAAAAAAA
BBBBBBBBBBBBBBBBBBBBBBBBB
Do you know how would be the best way to implement it?
thanks
• 709 views
•
link
0 answers
No answers yet.
Log in to answer this question.
You can try
clumpify.shfrom BBMap suite withcontainment=toption. Read more about clumpify here: Introducing Clumpify: Create 30% Smaller, Faster Gzipped Fastq Files. And remove duplicates. There is also a guide available.If this does not keep the longest representation of identical reads, you could filter your reads with
reformat.shorbbduk.sh(both from BBMap suite) with aminlength=option.