This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Non-Redundant Data Sets Of Protein

Hello everyone,

I'm working with plant proteome. As there are huge number of proteins in a proteome, I want to filter out the protein number within a limited length of amino acids. Is there any tool available for this? And anyone for calculating the average length of amino acid sequences in plant proteomes?

Thanks in advance.

What file type you're working with? Is it fasta?

Thanks. I've tried those. But it didn't work :(

1 answer

Jalview can do this fairly easily, though if you use an entire genome's worth of proteins it may run very slowly. Just sort by length through calculate, and then you can highlight the range of sequences you want to remove, and just hit delete. Then save the "alignment" as a new fasta file.

Is multiple sequence alignment required before using Jalview?

Log in to answer this question.