Hi all, I’m working on generating simulated metagenomic datasets to benchmark my binning pipeline. I want to simulate realistic community structures and sequencing errors. I need to have controllable ground truth for species and abundances, realistic sequencing error profiles (Illumina / long-read) and varying community complexity, that can be tailored by my requirements. I used ART simulation tool before but looking for something more convenient.
I have following questions: Which simulation tools (and parameters) produce the most biologically realistic and computationally useful metagenomic samples for benchmarking? How do people handle reference genome selection and abundance modeling (e.g., for soil vs gut microbiomes)? Any tips on simulating platform-specific noise (e.g., Illumina, Nanopore)?
I am considering using CAMISIM, so I would also appreciate any tips about it Thanks!
1 answer
See Mensur Dlakic answer here about pre-existing datasets that you can use (which have been produced for benchmarking) --> Metagenomics mock datasets
As for generating reads this is one option: Metagenome Read Simulators
Log in to answer this question.