Title: Different pathway predictions from KOfamScan/KEGG-Decoder and METABOLIC for the same MAGs
Hi everyone,
I am working with metagenome-assembled genomes (MAGs) and trying to reconstruct metabolic pathways. I am particularly interested in methanogenesis and methane oxidation.
I am getting substantially different results from KOfamScan + KEGG-Decoder compared with METABOLIC, even though both analyses were performed on the same MAGs.
My workflow is:
- Predict proteins with Prodigal.
- Run the predicted proteins through KOfamScan.
- Retain only KOfamScan hits marked with
*(hits passing the KO-specific threshold). - Use these KO assignments as input for KEGG-Decoder to evaluate pathway/module completeness.
In parallel, I ran the same bins through METABOLIC.
For some bins I obtain results such as:
- KEGG-Decoder: complete methanogenesis module (for example, M00357)
- METABOLIC: only partial module steps (for example, M00357+01), or the module is not reported as complete
In some cases, METABOLIC also detects individual methane-metabolism genes, for example mmoB or pmoABC, that are not represented in my KEGG-Decoder results for the same bins.
I suspect that some of the differences may arise because I am applying a strict *-hit filter to the KOfamScan results before running KEGG-Decoder, whereas METABOLIC uses multiple HMM resources and its own criteria for accepting functional hits. However, I would like to understand this more rigorously.
My questions:
- How exactly do the pathway/module completeness criteria differ between KEGG-Decoder and METABOLIC? I understand that METABOLIC uses a 75% completeness threshold for modules.
In particular, when KEGG-Decoder reports a module as complete but METABOLIC reports only partial steps, could this be caused by differences in how alternative KOs and module logic are handled?
- Does METABOLIC use functional evidence beyond KOfam/KEGG KOs when reconstructing these pathways?
If so, could this explain why METABOLIC detects some genes that are absent from my filtered KOfamScan dataset?
- Is retaining only KOfamScan
*hits an appropriate strategy for downstream KEGG module reconstruction?
I have seen some papers where additional filtering based on E-values was applied to KOfamScan results.
I understand that * indicates that the hit passed the KO-specific threshold, but I am wondering whether filtering only for * hits before pathway reconstruction could lead to false-negative pathway calls.
- For MAG-level metabolic reconstruction, what would be the most appropriate way to interpret these differences?
Would it be better to compare the underlying gene/KO assignments first, and only then compare the module/pathway predictions?
I am not trying to determine which program is universally "better." I would mainly like to understand the methodological reasons for the discrepancies and establish a defensible workflow for MAG-level metabolic pathway reconstruction.
For example, would a reasonable approach be:
gene/KO evidence -> module completeness -> biological interpretation
rather than relying on the final "complete/partial" pathway call from either program alone?
I thought that the same MAGs should give relatively similar predictions with different tools, so I am trying to understand what could explain these discrepancies.
If anyone has experience specifically comparing METABOLIC with KOfamScan/KEGG-Decoder for MAGs, I would greatly appreciate advice on how to interpret these differences.
Thank you!
0 answers
No answers yet.
Log in to answer this question.