This is a test version of Biostars. For the public version, visit https://www.biostars.org.
remove repeating sequences from multifasta file.

Hello,

im working with a database that contain some fasta files with my interest genes. But some FASTA files have sequences with different IDs with the same sequence.

So, i want to remove duplicate sequences based on the nucleotide sequence for make a nonredundant database

How i can do this?

thanks

gene genome

You are not very specific. One thing that is important to know here is how many sequences we are talking about. If it's say 500 then that probably fits in your computer's memory and we can eliminate these duplicates reasonably straight forward.

But if you are talking about millions of sequences then we need to come up with a more sophisticated strategy.

If two records with the same sequence are found, does it matter which one is deleted and which one is kept?

Hello savscosta!

Questions similar to yours can already be found at:

We have closed your question to allow us to keep similar content in the same thread.

If you disagree with this please tell us why in a reply below. We'll be happy to talk about it.

Cheers!

0 answers

No answers yet.

Log in to answer this question.