For the sake of my argument, the individuals are completely unrelated, i.e. related on the noise level. So the answers to 1-7 would all be something like "average" or "noise level" or "within one population". Then, after having an estimate of expected shared SNVs, I'd know the number should be higher for cousins or siblings.
It is a pairwise comparison, performed multiple times, so first you'd find SNVs shared by individual 1 and 2, then those with individual 3, etc (the number should decrease with every additional individual inclusion). An SNV is included on only one condition: its difference from hg38. I tried doing some basic compounding probability calculations but dropped it as it can't possibly be that simple. :/
see this response by Paolo Maccallini:
Thanks Jeremy, I am very impressed at how Xwitter posts can be embedded so easily...
Anyway, yes indeed that 4.5M would be the maximum - but what about the average? It should be much lower ja?
We could study this empirically on our 1000 genomes data. PM me and we'll discuss a scoring strategy (defining exactly what shared means)
Interesting! I'd love to discuss, but there's no PM function here, and you don't allow PMs on Xwitter. You can email me at joelwallenius at gmail, for example. :)