This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Add taxID to makeblastdb

I'm sorry if this is a repeated question, but I continue to have doubts.

I created a blastdb like this:

makeblastdb -in input.fasta -dbtype nucl -title test_DB -parse_seqids -out test_DB

But I can't understand how to add the taxid in order to have the same result has if I used the all nt database.

Can someone clarify me? Thanks

blastdb windows

Can you show what your input fasta headers look like? grep "^>" input.fasta | head -5.

1 answer

your command should look like:

makeblastdb -in input.fasta -dbtype nucl -title test_DB -parse_seqids -taxid_map taxidmapfile -out test_DB

The taxidmap file is a text file consisting of two columns. You can download the taxonomy id information here:

ftp://ftp.ncbi.nih.gov/pub/taxonomy/accession2taxid/nucl_gb.accession2taxid.gz

You need to unpack it and you can make a taxidmap file by doing (something like):

sed '1d' nucl_gb.accession2taxid | awk '{print $2" "$3}' > taxidmapfile

I did has you sad but then I received this error messagem: [makeblastdb] No sequences matched any of the taxids provided...

My fasta file looks like this: >NC_028405.1_COX1: I managed to solve this problem!Thanks for the help. But now I have another doubt, cause I finally managed to get the database that I want, but I got a different result and worst than when I used all NT. When I use a 'pre-made' database, shouldn't the results be better and more precise? Thanks for the help again.

When I use a 'pre-made' database, shouldn't the results be better and more precise?

In what way?

BLAST results are very dependent on search space (database size). This is vastly different between nt and any custom database you make so the results will be different (if you are looking at e values and such).

Log in to answer this question.