This is a test version of Biostars. For the public version, visit https://www.biostars.org.
search/subset multiple sequence alignment for all columns with specific ambiguity base (linux or seaview)

I have a multiple sequence alignment generated by MAFFT. I know there are some positions where one or more sequences have an ambiguity base other than N (e.g. K or Y or M etc.). I would like to subset my alignment across all species to just those residues/columns that have these non-N ambiguity characters in any of the aligned sequences. Any suggestions? As there are so few of these positions, I was even willing to do this by hand, using Seaview to identify the columns where any sequence has one of these bases (and then do the subsetting in linux), but the search function in Seaview only seems to search the selected row/sequence for a given string...not just the next occurrence of the string in the entire matrix...

alignment seaview sequence

0 answers

No answers yet.

Log in to answer this question.