This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Ideal coverage required for PacBio error correction using HGAP

What is the ideal PacBio coverage required for error correction (consensus polishing) using HGAP (by only using PacBio reads)? How effective is this approach in error correction?

How does this approach compare with hybrid error correction using Illumina short reads (assuming the reads are from a homozygous genome)?

pacbio hgap error-correction illumina

"issues for larger (and eukaryotic) genomes": related to assembly or error correction or both?

Please don't post comments as answers. Comment on the answer if you have a follow up.

1 answer

For HGAP only, recommended coverage would be 60x-100x (more repetitive genome, the higher). It seems to be very effective for small microbial genomes (you could end up with 1 or 2 contigs), it has issues for larger (and eukaryotic) genomes.

Comparison with hybrid approach will depend on many factors, most importantly genome size and complexity. For small microbial genomes, I wouldn't bother with Illumina and would do just PacBio and HGAP.

Thanks for the information. So far for microbial genome we don't need any error correction with Illumina reads. Do you have any idea after HGAP we can get many repeated contigs. How to deal with that? You have to remove by visually or any other approach?

Log in to answer this question.