This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Biopython: Automating Genpept Queries

Hi - I've been trying to use BioPython to automate some research I've been doing, and while I've been able to get the NCBIWWW module to download XML files for the queries I'm sending it, I'm not getting all the information I need. Specifically, I need the GenPept information (e.g. taxonomy of the organisms the query found) relating to each match. I can't find any way in the documentation to do this, though, so I'm turning to you for help. Can this be done?

Thanks for the help!

biopython genbank sequence retrieval

That's actually from the same project as what I'm on - it relates more to processing the GenPept files once we've got them. I'm trying to get them to download automatically in the first place. Thanks, though.

2 answers

You can access the NCBI taxonomy information using the NCBI Entrez Utilities. See the "Finding the lineage of an organism" example the chapter on Bio.Entrez in the Biopython Tutorial http://biopython.org/DIST/docs/tutorial/Tutorial.html and also http://eutils.ncbi.nlm.nih.gov/corehtml/query/static/efetchtax_help.html

Hi Ben, did you try to look at this ? ftp://ftp.ncifcrf.gov/pub/genpept/

There are 4 files here: gpdat_1.seq.gz gpdat_2.seq.gz gpdat_3.seq.gz gpdat_4.seq.gz

You can download and parse them accordingly, based on what you need. Hope it helps.

Log in to answer this question.