This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Illumina Hiseq insert sizes

Hello Everyone,

in the shotgun metagenomics field, to my knowledge Illumina Hiseq sequencing runs are virtually exclusively run with very similar insert sizes of very few hundred base pairs.

In regards to metagenome-assembled genomes (which - some people would argue - the field is inevitably moving towards), having such very small insert sizes seems to be sub-optimal: Assembly/scaffolding using Illumina HiSeq data with bigger insert sizes surely would work better.

Could anyone tell me why we are still using these small insert sizes? Is it easier to have consistently smaller fragments (as opposed to bigger fragments) when shearing DNA prior to sequencing?

Thanks, Nic

assembly

The main factor should be the read length. If you have 2x150bp then something with ISIZE a bit above 300bp, like 400-500 should work and everything beyond is probably not beneficial as it does not get sequenced anyway. Smaller fragments cluster more efficiently than longer ones, that is another factor.

(My previous answer got lost somehow, so here I go again)

If you have a DNA piece which is as long or longer than 2x150, changing the size of the DNA piece will not change the sequence information you generate (it will always be 2x150). What you do change, though, is the linking distance between the read pairs which should greatly benefit assemblers to span repetitive regions.

I don't understand your second point regarding smaller fragments clustering better. Which kind of clustering do you mean?

0 answers

No answers yet.

Log in to answer this question.