This is a test version of Biostars. For the public version, visit https://www.biostars.org.
clustering sequence FASTA GSS, EST, Transcripts

I want to do clustering (k-means) and redundancy removal of my FASTA sequences which are mainly GSS, EST and assembled transcripts, to create a reference set for my short query sequences. My short query sequences can target either DNA or RNA. So I need some expert guidance. Also should I convert lower case base sequences into upper case for doing this task. Any suggestion would be highly appreciated.

genome assembly sequence next-gen alignment

You should also look at CD-HIT which is specifically tailored for this type of application and has specific subprograms.

Thanks for your answer genomax, but I have found uclust to be better than CD-HIT

0 answers

No answers yet.

Log in to answer this question.