"Block quotes" in post editor used by OP was messing up the display. The headers are regular fasta type (see above).
Hi, I have a fasta file with 300 protein sequences. I intend to construct a phylogenetic tree with it. I would want only the accession number and the organism name in the fasta header and remove the rest of the information. Can anybody suggest how to do this? I have a linux based system with perl and python installed.
For example, i want to convert a header like this:
>gi|685204428|gb|AIN98665.1| fumarate hydratase, putative [Leishmania panamensis]
to a header like this
>Leishmania panamensis| AIN98665.1
Some sequences have multiple headers. Would that be a problem?
regards Vijay
2 answers
I ignored that part - my regex includes the > sign in the header. It's OP's statements on "multiple headers" that has me curious.
HI,
You could use biopython, something like (ive done no testing)
from Bio import SeqIO
fasta_sequences = SeqIO.parse(open(file),'fasta')
output_file="editedHeaders.fa"
myRecords=list()
for fasta in fasta_sequences:
originalID=fasta.id
newIDList=fasta.description.split('[')
species=newIDList[-1].replace(']','')
newID=species + "|" + newIDList[0].split('|')[3]
fasta.id=newID
fasta.description=""
myRecords.append(fasta)
SeqIO.write(myRecords, output_file, "fasta")
I would also add a dict to save the original headers and pickle it so you can look up if needed.
Good luck
Log in to answer this question.
Strictly speaking, yours is not a right FASTA. Anything following the first space/tab is not part of the sequence name. Renaming fasta like this may confuse other tools.