This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Identifying non-paralogous protein or nucleic acid sequences without using CD-HIT

Could someone assist me in identifying non-paralogous protein or nucleic acid sequences without relying on CD-HIT tools? Since this web server is no longer operational, what steps may we take to finish the task?

cd-hit

I think the other comments are valid/the right approach, but the open question here is why bother with a webserver at all? Just do a local install.

1 answer

You could use MMSeqs2 to cluster proteins in place of CD-HIT.

However, your question is pretty vague. Are you only using one species, or clustering proteins among many species. That changes which clusters and thresholds you would use.

EDIT: I am unaware of any web based protein clustering services. Can you get access to a unix machine? That would make things a lot easier.

Yes, I am working with one species and want to compare with human proteome.

I don't work on humans, but this seems like a resource that will already be well annotated somewhere. There is even a filter to exclude paralogous genes in biomart on Ensembl.

Log in to answer this question.