Hi,
I am annotating genome assemblies using Funannotate and having some issues. Specifically, I am working with Fusarium fujikuroi and having really frusturating issues, which may not be that big of a deal, but the program does not identify FUM21, an important transcription factor for the fumonisin biosynthetic pathway. I'm not sure how to manually edit this. I know the gene is there because all of my isolates are high fumonisin producers. Additionally, when I use RNA-seq data from IMI58289 (the reference isolate) as a training dataset, FUM21 pops up. However, this then introduces annotation errors that starts to split different genes in the FUM cluster and others into two or more separate genes. I have used Funannotate2, which does identify FUM21, but it combines it into a single large gene with the polyketide synthase gene, FUM1.
Overall, I'm not sure how best to move forward with this. I think the datasets that are lacking FUM21 detection may be the best to move forward with since that may be the 'simplest' to fix, although I do not know how.
I am reaching out hoping someone has some advice or insight that would be beneficial for moving forward. The FUM cluster is a particularly important one for my comparative genomics, and this makes it increasingly more difficult to work with.
Thank you.
1 answer
For funannotate, during the train, dont't just include RNASeq data from one reference. Atleast add 2/3 datasets of RNASEQ. You can add protein-evidences of close samples. like F.fujikuroi IMI58289 during the predictions stage along with funannotate database.
Still if this doesnot work, you can give BRAKER3 a try too
Log in to answer this question.
What are you using as your protein evidence for annotation without RNA-seq? I am guessing some sort of large protein dataset?
Are you able to add some copies of FUM21 to this in order to help with predictions?