This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Correcting homopolymer lengths with RNAseq data

Hi everyone

I'm working on some fungal genome assemblies based on Pacbio and Nanopore data and have run into the issue of varying homopolymer lengths that are causing frame-shifts in coding regions. I would like to correct/polish these regions with the existing data and/or with additional illumina RNAseq data that we have.

I was wondering what other people do to correct homopolymer length variation? Is there any other tools available other than Pilon that can utilise RNAseq data for polishing?

More specific details:

The genomes that we are working are ~ 50 Mb and we have ~20x coverage of Pacbio CLR and ~50x coverage of nanopore data (LSK109 on a R9.4.1 pore and basecalled using the high accuracy model in guppy 5.0.13). The genomes were assembled using >15 kb reads in canu 2.2. The Homopolymers are >=6 bp and length variation seems to occur more often in G/C homopolymers. I have tried using Pilon to polish the contigs using RNAseq data but I have read that Pilon doesn't polish homopolymers over 4 bp, in any case pilon isn't fixing the homopolymers. I've also tried tools like proovframe and Inspector or combinations thereof. I got the best results when combining all three of these but found Inspector trimming large regions (I still have to investigate this). It seems like polishing homopolymers using RNAseq data would be a good way to proceed but I haven't found any tools other than Pilon that can utilise RNAseq data?

Thanks for your help!

rna-seq polishing homopolymer pilon

0 answers

No answers yet.

Log in to answer this question.