Hi Pierre, I am doing something similar at the moment, and this looks like a very good solution. Maybe we should mention that the data is in genbank format (obviously) and needs to be converted to fasta before making a blastdb. However, when using BioPerl SeqIO to convert the fasta headers look like this:
>AB000048 Feline panleukopenia virus gene for nonstructural protein 1, complete cds, isolate: 483.
So, no gi's here, but they would be needed to assign taxids for metagenomics, any quick fix to keep the gi?