What version of kraken are your using? I cannot find this fille, I'm using the standard build database which produces these fills:
database.kdb database.idx taxonomy/nodes.dmp taxonomy/names.dmp:
I built a database from refSEQ in December, but my hits are missing a obvious genome.
Is there any easy way to extract a list of GIs used in the construction of a database or is this logged somewhere? Is there anyway to find out what subset of refseq kraken has built from?
seqid2taxid.map file in the Kraken DB directory contains the seqID and the taxonomy ID as two columns. If a database was built from scratch then the fasta files should be available in the library/added/*.fna in the Kraken DB folder.
What version of kraken are your using? I cannot find this fille, I'm using the standard build database which produces these fills:
database.kdb database.idx taxonomy/nodes.dmp taxonomy/names.dmp:
Log in to answer this question.
Did you find the way to extract the list of GIs from the Kraken DB? We are using as well one of the older versions of Kraken, and the only files we have are database files