This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Funannotate annotation help

Hi,

I am annotating genome assemblies using Funannotate and having some issues. Specifically, I am working with Fusarium fujikuroi and having really frusturating issues, which may not be that big of a deal, but the program does not identify FUM21, an important transcription factor for the fumonisin biosynthetic pathway. I'm not sure how to manually edit this. I know the gene is there because all of my isolates are high fumonisin producers. Additionally, when I use RNA-seq data from IMI58289 (the reference isolate) as a training dataset, FUM21 pops up. However, this then introduces annotation errors that starts to split different genes in the FUM cluster and others into two or more separate genes. I have used Funannotate2, which does identify FUM21, but it combines it into a single large gene with the polyketide synthase gene, FUM1.

Overall, I'm not sure how best to move forward with this. I think the datasets that are lacking FUM21 detection may be the best to move forward with since that may be the 'simplest' to fix, although I do not know how.

I am reaching out hoping someone has some advice or insight that would be beneficial for moving forward. The FUM cluster is a particularly important one for my comparative genomics, and this makes it increasingly more difficult to work with.

Thank you.

annotation genome

What are you using as your protein evidence for annotation without RNA-seq? I am guessing some sort of large protein dataset?

Are you able to add some copies of FUM21 to this in order to help with predictions?

1 answer

For funannotate, during the train, dont't just include RNASeq data from one reference. Atleast add 2/3 datasets of RNASEQ. You can add protein-evidences of close samples. like F.fujikuroi IMI58289 during the predictions stage along with funannotate database. Still if this doesnot work, you can give BRAKER3 a try too

Log in to answer this question.