This is a test version of Biostars. For the public version, visit https://www.biostars.org.
How to retrieve single protein fasta file for multiple species?

Hi all,

We are trying to make protein database of multiple organisms say E. coli, T. ferroxidans, B. subtilus, etc. This is what we want to use for matching our orbitrap output and we want to do that only with those species which we have found through Illumina sequencing. These are approximately 400+ genera. So, can you suggest any smart way of doing so? Like I provide the names of organisms and retrieve single fasta file?

Thank you very much!

fasta protein multiple_species database

You can use @5heikii's script here.

cating the individual fasta genome proteins files into a giant one afterwards should be a simple task.

Note: See new answer/commnet below.

running this code didn't generate any fasta file. Although both the list of species (species.txt) and assembly_summary.txt are is same folder. Am i missing something?

2 answers

Try this if you need RefSeq (modified version of @5heikki's code):

$ more species.txt 
Bifidobacterium adolescentis

$ wget ftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/assembly_summary_refseq.txt

$ IFS=$'\n'; for next in $(cat species.txt); do awk -v SPECIES=^"$next" 'BEGIN{FS="\t"}{if($8 ~ SPECIES){print $20}}' assembly_summary_refseq.txt | awk 'BEGIN{OFS=FS="/"}{print "wget "$0,$NF"_protein.faa.gz"}'; done | sh

Otherwise

$ wget ftp://ftp.ncbi.nlm.nih.gov/genomes/genbank/bacteria/assembly_summary.txt

 IFS=$'\n'; for next in $(cat species.txt); do awk -v SPECIES=^"$next" 'BEGIN{FS="\t"}{if($8 ~ SPECIES){print $20}}' assembly_summary.txt | awk 'BEGIN{OFS=FS="/"}{print "wget "$0,$NF"_protein.faa.gz"}'; done

You will get many strains etc by this method. If you need very specific strains then you could awk '{print $8,$9,$10}' assembly_summary.txt > species and only take those that you need.

thanks!! this worked perfectly. I have list of files such as GCF_000164035.1_ASM16403v1_protein.faa.gz and the next step would be to combine them together. Can you guide me there as well? Thanks a lot!!! :)

If you want the final data file uncompressed: zcat G*.gz > final.faa
If you want to keep the final data compressed: cat G*.gz > final.faa.gz

I have prepared another list of archea this time but this command is not working. Is there any other assembly summary for archea?

Post examples of names that are not working.

If you are working with UniProt, you can retrieve the data programmatically as described here (with code examples): https://www.uniprot.org/help/api_downloading https://www.uniprot.org/help/api_queries

Log in to answer this question.