+1, using correct tool and looks nice!
• 0 views
•
link
HI I am trying to split a fasta file using pyfasta, and print to individual files, but I'm having some trouble understanding the syntax. I am using:
pyfasta split --header afp2test.fasta afp2test.fasta
this runs, but doesnt split the files.
Wjat is the syntax i need?
GNU csplit is done for that kind of jobs:
csplit -z -q -n 4 -f sequence_ sequences.fasta /\>/ {*}
+1, using correct tool and looks nice!
this will do the job:
awk 'BEGIN {n_seq=0;} /^>/ {if(n_seq%1==0){file=sprintf("myseq%d.fa",n_seq);} print >> file; n_seq++; next;} { print >> file; }' < sequences.fa
it wil generate a file for ech of the sequences in your file 'sequences.fa'...
See https://pypi.python.org/pypi/pyfasta/
split the fasta file into one new file per header with “%(seqid)s” being filled into each filename.:
$ pyfasta split –header “%(seqid)s.fasta” original.fasta
You need to specify that sequence id (= the name of your fasta within multifasta file) will be used as the name of the new files. For that, seqid parameter is used.
Log in to answer this question.
I guess you need to say with what you wanna split the fasta:
extract sequence from the file. use the header flag to make a new fasta file. the args are a list of sequences to extract.
extract sequence from a file using a file containing the headers not wanted in the new file:
extract sequence from a fasta file with complex keys where we only want to lookup based on the part before the space.
but is there a way to do this without haveing to write down all the headers on the command line. i have a file with 100's of sequnces and i just want to make a new file for each sequence. the "split" command allows you to do this by specifiying the amount of files, but no matter what i do using this, the order does not seem to be preserved, and 1 file always contains 2 sequences. there is an option to split by header, but thought this would automatically pick up the individual headers and put them in their own files.