This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Identifying domain of long-read assembled contig

Hello all,

I have a metagenome with a whole bunch of assembled contigs. I'd like to pick out the bacterial contigs.

I first used Kaiju to classify these and identified ~20K bacterial contigs, but noticed many that were unclassified beyond the domain level were actually Eukaryotes based on Blast.

I then tried MEGAN6-LR, and identified 5K contigs. So far they seem more accurate, but there seems to be quite. big discrepancy and I fear I'm leaving a lot of data behind in false negatives using MEGAN.

Any tips?

longread taxonomy identification contig

you can consider running Kraken2 , or perhaps kma (but that is rather on read level )

Has that been shown to be more precise that both?

1 answer

In my experience, tiara works really well in classifying contigs between the 3 major kingdoms, plus mitochondria and plastids. The only requirement is that contigs are reasonably long, > 3000 bp. There is no need to install a large database of sequences, and the classification is fairly rapid.

Log in to answer this question.