I have used
grep -wEA1 --no-group-separator hsa-mir hairpin.fa > a.fa
which works fine with single-line FASTA sequences but I want to use it for multi-line FASTA sequences.
I have the entire hairpin sequences downloaded from the miRbase website. The first few lines look like this
cel-let-7 MI0000001 Caenorhabditis elegans let-7 stem-loop UACACUGUGGAUCCGGUGAGGUAGUAGGUUGUAUAGUUUGGAAUAUUACCACCGGUGAAC UAUGCAAUUUUCUACCUUACCGGAGACAGAACUCUUCGA cel-lin-4 MI0000002 Caenorhabditis elegans lin-4 stem-loop AUGCUUCCGGCCUGUUCCCUGAGACCUCAAGUGUGAGUGUACUAUUGAUGCUUCACACCU GGGCUCUCCGGGUACCAGGACGGUUUGAGCAGAU cel-mir-1 MI0000003 Caenorhabditis elegans miR-1 stem-loop AAAGUGACCGUACCGAGCUGCAUACUUCCUUACAUGCCCAUACUAUAUCAUAAAUGGAUA
I want to extract just the human hairpin sequences from the entire file. I have used the grep command as follows
grep hsa-mir hairpin.fa > human_hairpin.fa
However, it only extracts the header line but I need the sequences as well like this.
hsa-mir-548ab MI0016752 Homo sapiens miR-548ab stem-loop AUGUUGGUGCAAAAGUAAUUGUGGAUUUUGCUAUUACUUGUAUUUAUUUGUAAUGCAAAA CCCGCAAUUAGUUUUGCACCAACC
Which commands should I follow?
You obviously don't know it or else you wouldn't be asking, but your question is easily one of the top 3-5 most frequently asked.
Simply enter extract fasta sequences into the search box at the top and you will get many answers. For a single sequence, you can also open a file and copy-paste it by hand.
I have used
grep -wEA1 --no-group-separator hsa-mir hairpin.fa > a.fa
which works fine with single-line FASTA sequences but I want to use it for multi-line FASTA sequences.
Log in to answer this question.