Thank you!! That works for every possible flaw.
• 0 views
•
link
I have a FASTA file with numerous protein sequences (the header in one line and amino acid codes in several lines). Some of these sequences contain ambiguous or exceptional amino acid codes (e.g., B, J, O, U, Z, X, -- ). I want to remove sequences containing such code and generate a new FASTA file. How I can do this in python? I did manage the remove one amino acid code (X) at a time with the following code. But how can I remove them all at once?
from Bio import SeqIO
sequences = SeqIO.parse("sequences.fasta", "fasta")
filtered = [seq for seq in sequences if seq.seq.count('X') == 0]
with open('sequences_without_Xs', 'wt') as output:
SeqIO.write(filtered, output, 'fasta')
This code removes sequences that contain at least one character that is not an amino acid.
from Bio import SeqIO
AMINOACIDS = set('ACDEFGHIKLMNPRSTWVQY')
with open('sequences_valid.fasta', 'w') as output:
for seq_record in SeqIO.parse("sequences.fasta", "fasta"):
if not set(seq_record.seq).difference(AMINOACIDS):
output.write(seq_record.format('fasta'))
Thank you!! That works for every possible flaw.
Using the code below instead of your filtered line should do the trick.
filtered = [
seq
for seq in sequences
if seq.seq.count("X") == 0
and seq.seq.count("B") == 0
and seq.seq.count("J") == 0
and seq.seq.count("O") == 0
and seq.seq.count("U") == 0
and seq.seq.count("Z") == 0
]
Oh... Thank you!
Log in to answer this question.