It's not that simple. Even if proteomes are not complete, a protein in particular can have been already identified or characterized.
Dear all,
I would like to have the opinion of the community about a problem I'm facing. How to reconstruct phylogeny based on protein sequence of plant gene family. To this aim, one should retrieve all possible protein entries related to this family on Genbank.
Unfortunately as you probably know many of the protein sequences in GenBank (at the NCBI) are result of conceptual translations. Therefore they are predicted or hypothetical.
My aim is to infer the correct phylogeny without false positive/negative results, as well as not incurring mis-alignments due to incorrect predictions.
Which workflow/strategy would you recommend to choose ?
Thank you so much,
Luca
1 answer
You can choose plants for which the proteomes are complete (or at least close to it) in Swiss-Prot
Since many plants do not yet have complete proteomes, this would be somewhat limiting. As such finding as wide a range of family members as possible including taxa without complete proteomes it a reasonable thing to do in the first instance.
As a first pass searching UniProtKB/Swiss-Prot using either:
- Appropriate ontology and/or classification terms. For example: Gene Ontology (GO), InterPro, Pfam, etc. See UniProtKB cross-references for more options.
- Sequence similarity search. For example using BLAST or FASTA.
And limiting the result based on the Taxonomy annotations will give a set of possible candidates. This set can then be filtered based on the protein existence annotation. This will give a set of proteins that you can be reasonably sure actually exist in vivo. From there generating a phylogeny should be relatively simple.
Dear user, can I contact you in private? my email: ferrero64@hotmail.com
Log in to answer this question.
Look at ensembl plants for orthologs.