Hello folks,
I wonder if there are any scripts to extract 25% and 90% non-homologous data from protein databases like pdb or astral? I know this is a commonly done thing, so I was just thinking there had to be (openly available) scripts out there which people use regularly to get these datasets for their analyses.
thank you! Anj.
3 answers
Anj,
I think you should rephrase you question. Maybe I am missing the point but "25% and 90% non-homologous data from protein databases" does not make a lot of sense to me. First of all two things are homologous or they are not, we use similarity when we compare sequences with %s. Second % no-similar is still not understandable.
ok essentially i meant-how do i download the set of protein sequences which have only 25% sequence identity from a database? i know this is commonly done, and I was wondering if there are some commonly available scripts to do so?
--thanks!
If you are interested in non-homologous proteins whose structures are known, a widely used resource is the Astral dataset from http://scop.berkeley.edu/astral/
This dataset provides sequences (with known structure) with less than 40% or 95% identity.
Note that the 40% identity set will contain many homologous proteins. However, it is a useful first step.
Log in to answer this question.