Append number to the fasta sequence header with the times sequence has been repeated
Hi All, I need some help with programming this little problem (either in perl or python) which I am not very familiar with. Is there a way to append number to a duplicate sequence name(like Xtimes?) to the very first duplicate sequence and discard the other duplicates.
>name_1_1
AGGGTTT
>name_:2:_X
GTTTGAA
>name_:3:_Y
GTTTGAA
Result I want :
>name_1_1
AGGGTTT
>name_:2:_X_2times
GTTTGAA
• 2,252 views
•
link
2 answers
fasta are linearized. join two process: first process extract the DNA and count them. The second process sort the linearized fasta on the DNA.
join -t $'\t' -1 2 -2 2 <(awk '/^>/ {printf("%s%s\t",(N>0?"\n":""),$0);N++;next;} {printf("%s",$0);} END {printf("\n");}' jeter.fa | cut -f 2 | uniq -c | awk '{printf("%s\t%s\n",$1,$2);}' ) <(awk '/^>/ {printf("%s%s\t",(N>0?"\n":""),$0);N++;next;} {printf("%s",$0);} END {printf("\n");}' jeter.fa | sort -t $'\t' -u -k2,2) | awk '{printf("%s_%s\n%s\n",$3,$2,$1);}'
>name_1_1_1
AGGGTTT
>name_:2:_X_2
GTTTGAA
• 1 views
•
link
Dont know if you necessarily need to program it yourself but I use vsearch --derep_fulllength for this, easy and fast
• 1 views
•
link
Log in to answer this question.
mirDeep (Perl code) does exactly this. It is quite simple, just build a hash with sequences as key, and names as values. If a key already exist, instead of assigning the sequence name as value, append a counter.