This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Substitution matrix for Maximum likelihood tree and ASR

Hey, I'm new on this platform.

Me and my team want to do an ancestral sequence reconstruction focusing on a conserved domain, after having obtained the proteins, extracting the domain, curing them and running cd-hit. We have a total of 1265 sequences with an average lenght aroun 150 aminoacids with some being as long as 300 aa.

The important thing is that this conserved domains have a highly conserved E residue, followed by an EXXH motif after 25-30 aa. And it's symmetrical, so this are found twice in the protein when speaking from the whole superfamily.

Before building a maximum likelihood tree, we need to do the alignment of course, however, for what I've read MAFFT is usually recommended before using IQ-tree. However, this kind of alignment leaves too many gaps, and when wanting to change the parameters such as the subsitution matrix (We've been using Blossum 62), the iterations or gap penalties it takes forever to align.

We've been using thus Clustal Omega which has rendered a nice alignment with the desired conserved domains, however there are still too many gaps at the end and beginning of the sequence. When trimming using Trimal it delates a lot of evolutionary data which gives us an inaccurate tree (some subfamilies mixed with others and low bootstrap values).

I know the subfamilies beforehand, and we have added new proteins using blast and other tools. So when getting the tree we know more or less how it should and how it should not look.

So, here are my following questions:

  • Which type alignment would you recommend for my case and with which parameters?
  • How can I trim my alignments without losing too much information?
  • How long will these last?

Hope someone can help me, and please by patient as I'm a rookie.

asr alignment

0 answers

No answers yet.

Log in to answer this question.