thanks a lot mike!! it solved my problem!
Hello All,
I would like to sort the fasta header line (annotation). Below is the example of how my data is and it is in .txt
>AHF21055.1 ribosomal protein S4 (mitochondrion) [Helianthus annuus]
>AAM96597.1 ATP synthase F0 subunit 6 (mitochondrion) [Chaetosphaeridium globosum]
>AAM96598.1 ATP synthase F0 subunit 8 (mitochondrion) [Chaetosphaeridium globosum]
>AAM96599.1 ATP synthase F0 subunit 9 (mitochondrion) [Chaetosphaeridium globosum]
I would like to get the data as below: just the accession number and protein name preferably in table format and remove everything after the protein name.
example:
>AHF21055.1 ribosomal protein S4
>AAM96597.1 ATP synthase F0 subunit 6
>AAM96598.1 ATP synthase F0 subunit 8
>AAM96599.1 ATP synthase F0 subunit 9
Thank you in advance!
2 answers
This gives me the exact output you want as long as (mitochondrion) is present in all lines:
cat old_fasta_headers | sed '/^[[:space:]]*$/d' | cut -d\( -f1 | sed 's/\(\.[[:digit:]]*\) /\1\t/g ; s/$/\n/g' \
> new_fasta_headers
Hope this helps.
From my experience, FASTA headers consist of two parts - the ID and the description. You can use a tool like bioawk to extract just the
identifier and then sort the output, or you can use any combination of command line utilities, such as grep -o or cut or sed, much like Eric Lim's comment.
Log in to answer this question.
What have you tried?
PS: Please use the formatting bar (especially the
codeoption) to present your post better. I've done it for you this time.Sure Ram will do that from next time. Thanks a lot! I am kinda new to this forum
Assuming the (mitochondrion) is always there, this is what I can think on the of my head
cut -f1 -d'(' header.txt | sort. There will be an empty space at the end and can be removed bysed 's/ *$//'.thank you Eric Lim for your reply!
Do all of your entries follow that format? Will there be some where the string
(mitochondrion)is not there?