This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Feedback on Workflow

Hello

I'm a PhD Student doing WGS for the first time and I was wondering if anyone could possibly offer some feedback on the pipeline I've put together. I have Paired-end Data which I would like to make the most of.

I'm not sure if I've made some glaring omissions or if any of these steps are redundant.

Any feedback would be greatly appreciated

enter image description here

wgs pipeline

what's the end goal? Is this metagenomics study?

Sorry I didn't mention that.

It's a metagenomics study, I'm looking to compare "Healthy" and dysbiotic microbiomes. The goal is to do Taxonomic and Functional Analysis.

In that case I would suggest biobakery pipelines as well.

Thank you, the HUMAnN pipeline looks really useful It looks like it does everything I'm interested in.

It looks like the pipeline takes only 1 file as input per sample, Are there any steps in my pre-processing I may have overlooked when trying to consolidate my data into a single fastq file.

1 answer

There are already pipelines in place for everything these days. If this is your first time working with NGS and the sequencing is pretty standard I would suggest using those established pipelines. Take a look at nf-core

Thanks, I'll check out nf-core. Having a quick look there doesn't seem to be anything for analyzing Metagenomes with MEGAN but there looks like some interesting things I'm definitely interested in learning other tools down the line.

I think I was really treating it as a learning exercise to make sure I understood how all the steps of a pipeline fit together. Ideally, I'd also like to use my R1 and R2 data, and DIAMOND or is it meganizer only accepts one file as input.

Log in to answer this question.