This is a test version of Biostars. For the public version, visit https://www.biostars.org.
How to cluster existing multiple sequence alignments to identify homologous clusters

I have a number of existing multiple sequence nucleotide alignments from closely related taxa (two clades which are sisters), and need to align these alignments for analysis. Some are homologous and some not. I think the best way is to cluster them all together to identify these homologous clusters. I know how to do this for single sequences but not entire alignments.

clustering sequence fasta multiple alignment nucleotide

How about just pooling all the sequences and clustering with e.g. cd-hit or vsearch?

Hi, I've ran cd-hit to identify clusters and aligned them, but some are fragment sequences with full counterparts that need to be merged, and there are too many to go manually. So I'm looking to identify the pairs of clusters likely to be homologous and align them

Hi, thanks yeah I have mafft in mind for the aligning task but before that I'd like to cluster homologous pairs of alignments, because some are made of fragmentary sequences with full-sequence counterparts.

I see, perhaps you can first run a simple blast and run blastclust or such on the results to get a rough clustering ?

0 answers

No answers yet.

Log in to answer this question.