Hi Philipp,
Thanks for your feedback. I found that 0.1 outputs includes most of the species the authors claim to be in their mock dataset but it also outputs some that are not supposed to be there. On the other hand 0.2 give less taxons with less correct calls, in both cases I get Klebsiella. I understand the problem you mention but there are two things that I can't fully wrap my head around:
1) According to the authors there should be no enterobacteria in the dataset, the species listed there would be better classified as other taxa in the absence of true hits in the database. 2) I've used this database with another approach (dada2) on 16S sequences and it outputs proper classification of multiple enterobacteria, I don't see why in this case I'm having false Klebsiella false positives on very different samples.
The database is supposed to be pretty well curated and complete (I'm talking about the SILVA 138 Ref NR99). I'm running the analysis again using a --confidence of 0.15 to see if something better comes out.