thanks! I'll look through that post... of course I'm probably being a little too nitpicky with my search... we'll never really find ideal data in the real world, will we?
I'm looking for either a database (like SRA) or even a study that provides its data that has labels associated with the data. Ideally, this would be metagenomic data (either sequences or abundance tables) in a study that has a strong link between a feature like a species and the condition being studied.
Just reaching out because I haven't been able to find any studies that have enough data for my application (implementing machine learning algorithms) - so ideally we are talking about at least 100 samples for the condition being studied (controls, maybe the same).
Any help is appreciated. Thanks guys
2 answers
Hi Edward,
You can check below mentioned post. It may solve your purpose http://github.com/gjospin/PhyloSift/issues/59.
The closest databases I can think of are
- Repositive (https://discover.repositive.io/), which is more general purpose sequencing data
- Qiita (https://qiita.ucsd.edu/), which is microbial sequencing specific
- American Gut Project (http://americangut.org/), where some data can be found at https://github.com/biocore/American-Gut
The American Gut project has the quantity of data, so that might be the most interest to you for machine learning purposes. Good luck.
Log in to answer this question.