Thank you so much. Got it. I highly appreciate your help and support.
which is better, if
- query coverage is 100% and percentage identity of 69%, Max score: 675
- query coverage is 100% and perc. identity is 71% Max score: 670
- query coverage is 99%% and perc. identity is 69%, Ma score: 679
- query coverage is 98% and perc. identity is 72%, Max score: 674
I have to find the nearest bacterial species (homolog) to the Human fumarase protein sequence. I have got the above results. how can I select the one nearest bacterial species? I have to choose only which is the nearest to the human protein sequence I have put as a reference
1 answer
For practical purposes, there is no difference between these results. You would not do wrong by picking any of them. You can see that your #2 candidate has the lowest score, even though it has the highest coverage and the second highest identity. Based on those two features alone one would expect that particular sequence to have the highest score. Yet it must not be as close as others to the human protein in non-identical portions of the sequence, which brings down its score. Still, picking #2 is easy to defend, because there are very few known examples of protein pairs that are of about the same length and >70% identical that are not functional homologs.
If you know anything about catalytic residues, it would be best to pick the homolog that has those residues identical to the human protein. My guess is that all of them will be equal in that regard.
Log in to answer this question.
If you are referring to 4 different results/strains then all of them can be that candidate. You will want to evaluate if critical amino acids are conserved or not before making a decision.
Yes, I am referring to 4 different bacterial strains. Actually, I have to select only one which is the most nearest. shall I choose No:2 as it represents 100% of query coverage and 71% of identity even though it has a lower max score value, but others got less than 100% of query coverage?