Thank you very much!
• 0 views
•
link
I got a bunch of genome sequences in the same fie named sequence.fasta but some of them have the exact same names, like this:
> Rhodobacter_sphaeroides_2.4.1_chromosome_2
ATGAGCTTTCCCCATTTCGCGGCCCTCTTCCGGCCCTCGCAGTTCTTCGGCATCCGCGGCGGCGTCCACCCCGAGACGCG
>Rhodobacter_sphaeroides_2.4.1_chromosome_2
GTGCAGGTGGTGCCGACCCAGTATCCGATGGGCTCGGAGAAGCATCTGGTGAAGATCCTGACCGGGCGCGAGACGCCGGC
Is there any way to detect those sequences with the same name and add suffix automatically, so i can distinguish. this is what i want:
> Rhodobacter_sphaeroides_2.4.1_chromosome_2.1
ATGAGCTTTCCCCATTTCGCGGCCCTCTTCCGGCCCTCGCAGTTCTTCGGCATCCGCGGCGGCGTCCACCCCGAGACGCG
> Rhodobacter_sphaeroides_2.4.1_chromosome_2.2
GTGCAGGTGGTGCCGACCCAGTATCCGATGGGCTCGGAGAAGCATCTGGTGAAGATCCTGACCGGGCGCGAGACGCCGGC
But for those who have unique names just leave them.
Thanks a lot!
linearize, sort, count the uniq names:
awk '/^>/ {printf("%s%s\t",(N>0?"\n":""),$0);N++;next;} {printf("%s",$0);} END {printf("\n");}' input.fa | sort -t $'\t' -k1,1 | awk -F '\t' 'BEGIN{N=0;prev="";}{if(prev==$1) { N++;} else {N=1;} printf("%s.%d\n%s\n",$1,N,$2);prev=$1;}'
>1_anotherUniqueGeneName.1
atgc
>1_duplicateName.1
atgc
>1_duplicateName.2
atgc
>1_uniqueGeneName.1
atgc
Thank you very much!
With BBMap's reformat.sh:
reformat.sh in=file.fa out=fixed.fa uniquenames
That appends "_2", "_3", etc to the second and 3rd instance of a name. The first time a name occurs it will be unaffected.
Log in to answer this question.
before the name there is a ">" so it's like this
http://bioinf.shenwei.me/seqkit/usage/#rename