This is a test version of Biostars. For the public version, visit https://www.biostars.org.
pooling reads for error correction

For error-correction of Illumina reads, is it a good idea to pool reads from several samples and do error correction on the pool, rather than individually? (The samples are from patients with the mumps virus.) The intuition is that reads in region of low coverage might look like errors when seen in isolation, but would be confirmed by similar reads from other samples.

Are there error correction tools that explicitly take several samples, and give more credence to kmers/substrings seen in several samples?

rna-seq assembly error correction preprocessing

1 answer

I don't think it is a good idea to pool samples for error correction, and I also don't think it is a good idea to pool samples to assemble. What are you trying to assemble anyway? I guess it is not the human transcriptome. Are you trying to assemble the virus genome?

Hey there,

I am also interested in this question. Care to expand on why you don't think its a good idea to pool samples to error correct or assemble? I have read a few papers describing metagenomic assemblies where they preformed "co-assemblies", as in pooled all their reads and assembled.

thanks.

Log in to answer this question.