You're right, I wasn't explaining the problem clearly. Thanks for the directions!
- The depth; coverage ~3.00X ± 2.50 SD
- The sequencing platform; NovaSeq 6000 S4 (Illumina, CA)
- The preprocessing steps employed; https://bitbucket.org/wegmannlab/atlas-pipeline/wiki/Home with Atlas
- The aligner; BWA
- The variant-caller; https://bitbucket.org/wegmannlab/atlas-pipeline/wiki/Home with Atlas
- The variant-caller's parameters; Have to check ... There was an MAF of 0.05.
- Whether the DNA was amplified, and how much; PCR, 13 cycles
- but other manipulations were done because of bubbles Whether you are
- talking about individual or population studies; These are individuals
- And a legend for graph. red = AltAlt, blue = RefAlt, green = RefRef
The graph is based on per locus allele frequency (x axis) in relation to the 3 genotypes frequencies possible at each loci (red = AltAlt, blue = RefAlt, green = RefRef). This graph is basically showing the expected genotype frequencies (Hardy-Weinberg; HW, as the lines in the plot), with the 'real' genotype frequencies found in the population for each loci. This is way off what is expected and almost all sites are significantly different from what is expected assuming HW. I expect to find some sites that are not in HW but not all of them! (see something like this: https://gcbias.org/2011/10/13/population-genetics-course-resources-hardy-weinberg-eq/, note that to get AF>0.5, they randomly selected sites to be on the other side, but you are right, I don't know on top of my head, why there is no AF >0.6 here...)
The thing is that we looked at lcWGS data that was already published and found the same excess heterozygosity.