This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Kmer Frequency Distribution And Genome Complexity

I have a question regarding the kmer frequency distribution and Genome complexity.

I have the kmer distribution numbers (from kmergenie) for a newly sequenced genome of lizards, which are known to be highly polymorphic and have a highly repetitive genome. The distribution shows peaks at different kmer frequency numbers. For instance 19, 25, 29, 37, 49, 59, 69, 79, 83, 89. I do not understand this multiple peaks and multiple sub-peaks, including sub-peaks at half the kmer of a bigger peak.

I have read that kmer frequency peak at lesser than 20, about 17 means bacterial contamination. And a sub peak at half the value of a main peak means polymorphism.

http://arxiv.org/pdf/1308.2012.pdf - Heterozygosity and halk peak

https://groups.google.com/forum/#!topic/bgi-soap/xKS39Nz4SCE - Polymorphism with multiple peaks

But I cannot make anything out of the pattern I have right now. Does it mean there are chances of contamination. Or it that the genome is heterozygous, polymorphic and repetitive genome? Or something that I am missing??

ngs genome

I disagree with about 17 being bacterial contamination. It really depends on how much contamination you got. If you have large bacterial contamination, the peak will be higher, if you have lower bacterial contamination, it will be lower.

The lizard is diploid, so you are to expect 2 peaks, one representing homozygous regions and one peak that is exactly 50% the kmer coverage of the homozygous peak, the second peak is heterozygous variation.

Everything else can be explained by contamination. Can you show the graph, preferably at k=17?

kmer range

k18

k16

k genomic.kmers 14 29669462; 16 4651501; 18 5626211; 20 5826052; 22 5272846; 24 4555244; 26 4560471; 28 4162206; 30 4227152; 32 3835572; 34 3833445; 36 3810373; 38 3821880; 40 3648982; 42 3501479; 44 3399287; 46 2784813; 48 2474182; 50 3088075; 52 2979740; 54 2824460; 56 2757821; 58 2345110; 60 2654061; 62 2587238; 64 2507417; 66 2215479; 68 2118323; 70 2236174; 72 2100006; 74 2011707; 76 1892799; 78 1499840; 80 1697554; 82 918819; 84 1422869; 86 1225194; 88 614925; 90 740498

Sorry for the bad format of the frequencies, hope it is readable.

P.S. Open image in a new tab

Could you upload the Kmergenie HTML report file somewhere? It isn't clear that k=116 would be the best one to look at.

What the approximate genome size you are expecting (10 Mbp, 100 Mbp?), and estimated sequencing coverage of your data?

The approximate Genome size is 1.8-2.0 Gb. And the coverage depth is about 10X.

As you can see, the predicted assembly size is 5Mb, far less than your estimate of 1.8Gb. Also, at 10X genome coverage, there really isn't much to do with kmers.

Ok... But when I tried with Minia I got a genome size of >300Mb. And I think if I try Soapdenovo or Velvet I would be able to get 2GB. This happened with some other dataset I had before.

with velvet, what kmer did you use?

I did not try velvet or soap with this dataset. But I did try the Minia assembler with a kmer of 27 and 29, each of which gave contigs of >390Mb

Thanks for the HTML file. I agree with akoik063, 10x coverage is not sufficient. Minia is able to assemble some of the regions but the assembly will likely be of terrible contiguity and coverage. The same will be true with SOAPdenovo or Velvet.

Regarding Kmergenie, it predicts a tiny genome size because it thinks that the low-coverage k-mers are erroneous (and doesn't count them in the genome size).

To get a better assembly, you should have at least 30x coverage with Illumina data.

0 answers

No answers yet.

Log in to answer this question.