Could this be due to the starting DNA concentration (wood is a little hard to get better yield than other samples, particularly when wood is not completely decomposed)?
I am singling out this question, but this is an answer to just about all your questions. Binning is done with no regard to how the data is collected, what types of organisms are there. The only thing that matters is that there are contigs of sufficiently large size (larger than 1000 nucleotides for MetaBAT2, don't know about Maxbin2) and that these contigs have a tetranucleotide signal that clearly separates them into groups. You can rationalize that through sample collecting any way you want, but for our purposes it is the quality of your assembly that is relevant rather than your collection method. We don't know anything about the former, for example the number of contigs, average contig size, the N50 value, etc.
With the information you provided, an educated guess is that some of your assemblies are of poor quality - likely fragmented.
Should I use contigs.fasta instead of scaffolds.fasta?
It should make no difference, as in most cases those two are identical. Again, if you do the statistics on both files, you will probably arrive at the same conclusion.