Thank you for this, Kevin. We are also trying to wrap our minds around the best diversity descriptive statistics to provide for high-throughput sequencing data. Observed and expected heterozygosity have been staple descriptive statistics for microsatellite data sets for a long time, which may be why some are keen to figure this one out!
With that said, I'm having trouble understanding the formula on observed heterozygosity above. Although I'm still pretty new to this, wouldn't the following formula be more indicative of proportion of observed heterozygosity?
N_Sites - O(HOM) = O(HET)
N_Sites - E(HOM) = E(HET)
Then, after that, you can do the following formula to provide an individual proportion of expected and observed heterozygous sites:
O(HET) / N_Sites = Proportion Observed Heterozygous Sites
E(HET) / N_Sites = Proportion Expected Heterozygous Sites
All thoughts welcome if I'm far off base here.
Dear All,
Is there any explanation for the above please?
Thanks,
Ahmed
I am having the same question. Can somebody answer here. I did the same using --het command of plink and get the same output mentioned above.What next should i do to reach to observed and expected heterozygosity
Did you find an answer? I'm stuck on this too
The answer is below