Please upvote the script provider in that question
• 0 views
•
link
Hello,
How can I export some features (like country and date) from sequences inside a genbank file?
Thank you in advance!
Please upvote the script provider in that question
The " LOCUS" tag typically contains the date at the end like mentioned below:
LOCUS NM_001204686 968 bp mRNA linear INV 28-MAY-2017
This could be extracted like this:
grep "LOCUS" <your_genbank_file.txt> | awk '{print $NF}'
I am not sure what do you mean by "country" information. Please elaborate or provide example.
I mean from where the sequences come. Like in this exemple:
LOCUS KT968663 8103 bp RNA linear VRL 29-MAR-2016
DEFINITION Foot-and-mouth disease virus - type A isolate A/HY/CHA/2013,
complete genome.
ACCESSION KT968663
VERSION KT968663.1
KEYWORDS . SOURCE Foot-and-mouth disease virus - type A (FMDV-A)
ORGANISM Foot-and-mouth disease virus - type A
Viruses; ssRNA viruses; ssRNA positive-strand viruses, no DNA
stage; Picornavirales; Picornaviridae; Aphthovirus.
REFERENCE 1 (bases 1 to 8103)
AUTHORS Yang,X., Yang,J., Wang,H. and Zeng,F.
TITLE Direct Submission
JOURNAL Submitted (30-OCT-2015) College of Life Science, Sichuan
University, Wangjiang Road 29#, Chengdu, Sichuan 610064, China
COMMENT ##Assembly-Data-START##
Sequencing Technology :: Sanger dideoxy sequencing
##Assembly-Data-END##
FEATURES Location/Qualifiers
source 1..8103
/organism="Foot-and-mouth disease virus - type A"
/mol_type="genomic RNA"
/serotype="A"
/isolate="A/HY/CHA/2013"
/host="Bos grunniens"
/db_xref="taxon:12111"
**/country="China"**
/collection_date="15-Aug-2013"
/note="subtype: Sea97"
in a shell file, say "run.sh"
grep "^LOCUS" gb.txt | awk '{print $NF}'
grep "country" gb.txt | sed 's/ //g' | cut -d "=" -f2 | sed 's/"//g'
run as
$sh run.sh
output
[user@desktop]$ sh run.sh
29-MAR-2016
China
Log in to answer this question.
Did you go through this biostar question?
you can use this python program to extract different fields from gene bank file https://github.com/dewshr/NCBI-Genbank-file-parser