Does your method take into account Pfam Clans (which group together related families?). In my quick scan of your paper, I couldn't see that you'd done that..
What would be a painless way to get a collection (100 or 1000 etc) of amino acid sequences of proteins with only one protein per family? I am terrible at navigating databases for this, so if anyone has a step-by-step solution (or even better, knows of such a list that is already made that I could download), I would be so, so, grateful!! :)
Thank you
2 answers
Sorry for shameless plug :) - I have worked on characterizing best-respresentative sequence from protein families.
Please see: 3PFDB database - a database of best-representative PSSMs (and sequences) derived from protein families using an . Manuscript is available here.
Also check this short conference report on gathering best representative sequences.
Please let me know if you specifically need any datadumps or any additional infomraiton.
I remember, I did this 7 yrs ago: Go to pfam, look for Pfam-A seed sequences used to buld the pfam family profile and select one with longest length.
Rm thanks for your reply! I do not have any specific protein, though, I would like to generate a list of proteins (any kind of proteins) so long as they are not from the same family. Is what you are saying mean I would need to search for a specific protein? Because I am not searching for a representative from any specific protein(s). Do you see what I mean?
Log in to answer this question.
What do you mean by family? A sequence-based classification (e.g. PFAM) or structure-based (e.g. SCOP)?
Wow, I did not see your comment until now! I did not even realize there were different classifications. I just do not want to use proteins that may be functionally similar or gene duplicates. I am using HMM on the data, and want to mitigate any dependency. Any idea? And thanks!