This is great venu - thank you :)
• 1 views
•
link
Hi, I get my sequence files like this
TAACAGTGGGTCGAGATAAGGAC 1
CGCCCGGACGTCAGAGAAAAAGGTGCCAGCGCGGCGCAGAGAATGGAATT 1
TTAGTATGACGGGCTGACACGAGA 3
I want to add a sequence name at each sequence file as like:
>seq000001 1
TAACAGTGGGTCGAGATAAGGAC
>seq000002 1
CGCCCGGACGTCAGAGAAAAAGGTGCCAGCGCGGCGCAGAGAATGGAATT
>seq000003 3
Thanks, Fuyou
Assuming every sequence is of one line and you need to create the sequence headers
wc -l sequences.fa
Lets say you've got 100
for i in {1..100}; do echo "Seq000"$i >> ids.txt; done
paste ids.txt sequences.fa | awk '{print $1"-"$3"\t"$2}' | tr '\t' '\n' | sed 's/^/>/;n' | sed 's/-/ /g' > new_file.fa
Output
>seq0001 1
CAGTGGGTCGAGATAAGGAC
>seq0002 1
CGCCCGGACGTCAGAGAAAAAGGTGCCAGCGCGGCGCAGAGAATGGAATT
>seq0003 3
TTAGTATGACGGGCTGACACGAGA
This is great venu - thank you :)
Playing awk-golf:
awk '{printf ">seq%06d %d\n%s\n" , NR,$2,$1}' sequences.fa > new_sequence.fa
This takes the line number (NR) as the number for the seq-id, paste the last number in the fasta-header, and prints the sequence in a separate line.
All under the assumption, that you have per line one sequence followed by the number.
Log in to answer this question.
You seem to enjoy using custom-made formats :P
Where does the sequence name data come from Fuyou?
Hi John, Thanks, The seq name only need randomly name. It like I have wrote name, seq000001,seq000002,seq000003. Thanks