This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Different pathway predictions from KOfamScan/KEGG-Decoder and METABOLIC for the same MAGs

Title: Different pathway predictions from KOfamScan/KEGG-Decoder and METABOLIC for the same MAGs

Hi everyone,

I am working with metagenome-assembled genomes (MAGs) and trying to reconstruct metabolic pathways. I am particularly interested in methanogenesis and methane oxidation.

I am getting substantially different results from KOfamScan + KEGG-Decoder compared with METABOLIC, even though both analyses were performed on the same MAGs.

My workflow is:

  1. Predict proteins with Prodigal.
  2. Run the predicted proteins through KOfamScan.
  3. Retain only KOfamScan hits marked with * (hits passing the KO-specific threshold).
  4. Use these KO assignments as input for KEGG-Decoder to evaluate pathway/module completeness.

In parallel, I ran the same bins through METABOLIC.

For some bins I obtain results such as:

  • KEGG-Decoder: complete methanogenesis module (for example, M00357)
  • METABOLIC: only partial module steps (for example, M00357+01), or the module is not reported as complete

In some cases, METABOLIC also detects individual methane-metabolism genes, for example mmoB or pmoABC, that are not represented in my KEGG-Decoder results for the same bins.

I suspect that some of the differences may arise because I am applying a strict *-hit filter to the KOfamScan results before running KEGG-Decoder, whereas METABOLIC uses multiple HMM resources and its own criteria for accepting functional hits. However, I would like to understand this more rigorously.

My questions:

  1. How exactly do the pathway/module completeness criteria differ between KEGG-Decoder and METABOLIC? I understand that METABOLIC uses a 75% completeness threshold for modules.

In particular, when KEGG-Decoder reports a module as complete but METABOLIC reports only partial steps, could this be caused by differences in how alternative KOs and module logic are handled?

  1. Does METABOLIC use functional evidence beyond KOfam/KEGG KOs when reconstructing these pathways?

If so, could this explain why METABOLIC detects some genes that are absent from my filtered KOfamScan dataset?

  1. Is retaining only KOfamScan * hits an appropriate strategy for downstream KEGG module reconstruction?

I have seen some papers where additional filtering based on E-values was applied to KOfamScan results.

I understand that * indicates that the hit passed the KO-specific threshold, but I am wondering whether filtering only for * hits before pathway reconstruction could lead to false-negative pathway calls.

  1. For MAG-level metabolic reconstruction, what would be the most appropriate way to interpret these differences?

Would it be better to compare the underlying gene/KO assignments first, and only then compare the module/pathway predictions?

I am not trying to determine which program is universally "better." I would mainly like to understand the methodological reasons for the discrepancies and establish a defensible workflow for MAG-level metabolic pathway reconstruction.

For example, would a reasonable approach be:

gene/KO evidence -> module completeness -> biological interpretation

rather than relying on the final "complete/partial" pathway call from either program alone?

I thought that the same MAGs should give relatively similar predictions with different tools, so I am trying to understand what could explain these discrepancies.

If anyone has experience specifically comparing METABOLIC with KOfamScan/KEGG-Decoder for MAGs, I would greatly appreciate advice on how to interpret these differences.

Thank you!

data shotgun

0 answers

No answers yet.

Log in to answer this question.