Thank you for your honest assessment. I believe you have added much value to this discussion. As part of the lay, I was hoping for a good meaty answer.
To the laity, your discussion is quite disturbing. Consider a more direct question, with the information you provide, one could ask: can we create a useful WGS at all? This is disturbing because the medical industry is trying to use WGS data while its basic structures and patterns are still being developed. With that said, I would argue that we are learning a lot via this discussion. So let's explore a few more basic questions at play, beyond the Y issues raised above. This might teach us more about the basics quickly and give us a better understanding of WGS current state then the advertised sales speech of genetics industry.
Let's for these discussions, assume for the moment one can map a useful WGS. If one looks statistically across a WGS, what should the read distributions be? A reference in this case is meaningless. We are interested in the distribution of reads -- this really looks at the base quality of the reading process given that we can read a single chromosome pattern or a dual chromosome pattern. In this case, one is only really interested in homogeneous reads verse non homogeneous reads. I would argue, Fig. 4A and B would be a reasonable expectation for a diploid system. Moreover, Fig 4 C, one could argue, would be more expected for a haploid system with a high quality reading process.
Perhaps your comments are well taken for Figure 4E and one can't know what to expect yet because Y is different than diploid or haploid. Is that a fair over simplification of your discussion? I base this over simplification on an understanding that your general argument suggests there are far greater levels of mutation and other processes changing the Y chromosome overtime? So numerous variants can exist all at once for many reasons.
Nevertheless, using any sampling of chromosomes within WGS data, should there exist a typical patterns of the reads? Perhaps one can't discuss what to do with Y yet because it doesn't fit into this simple two bin system.
But within a diploid or haploid assumption, do Figs 4C vs 4D(4A/B might be better) form a reasonable pattern useful to differentiate a diploid from haploid systems, regardless of a reference. This argument is not based on a reference notion but merely a statistical notion. Mainly, your are either sampling one chromosome or two along the sequence. So the read should either be the same or split into two results at each location. Moreover, does it seem reasonable for a mosaic system to be some blend in between Fig. 4C and 4D, irrespective of the "true reference sequence." So perhaps the question of X verse XX can still be rightfully examined ill-respective of the reference sequence, simply considering raw differences of each chromosome read distributions. (I believe we also need to assume one common X within the two systems.)
Given these thoughts, what changes should be expected within Figs. 4 A,B,D? How would one expect the .5 hump to be transformed by a mosaic set of chromosomes made up of mixed diploid and haploid cells? Consider for argument 90% to 10% respectively. How should the .5 hump shift in probability and change its height based on the percentage of mosaicism? Moreover, what would one expect a mosaic system's read distribution pattern to look like? I think those are useful thought questions.
From a lay perspective, I think these are relative questions in helping one's deeper understanding of WGS datasets at their most basic levels with current technology and explained science. Also, there needs to be some type of test that flags questionable WGS results when basic assumptions are violated like additions of mosaicism are introduced within a WGS.
doi: 10.1093/gigascience/giz074
So I am including a journal article related to the question, I believe, and an additional tool that can be used to evaluate the issue. The tool is call xyalign. It is a command line boiconda tool. The journal and article come from is around a 7.5 impact rating, which is suppose to be a very highly rated journal. This should imply the information is more reliable than most publications. Based on the paper's Fig. 4E, most Y chromosomes are some what heterozygous within WGS datasets. Fig. 4E shows an example distribution of heterozygous to homozygous reads within a sampled WGS. This directly implies Y chromosome are to some extent heterozygous. Thus explaining what I have seen in the data (perhaps). Therefore, this should be expected within the general population.
Perhaps a formal geneticist can help confirm this notion.
Furthermore, the xyalign tool can be used to create the similar Fig. 4 plots for any WGS dataset. I believe that the plots in Fig. 4C and 4D can be used to infer intersex mosaicism. The relative allele peak around 50% read line might imply mosaic percentage of the WGS. If geneticist's would like to discuss this more that would be great! I think this kind of test can help validate and answer WGS dataset questions in the future. Furthermore, the relative height of the peak within Fig. 4E might also imply loss of Y chromosome percentages. This is a hot genetics topic in the world today. Perhaps there are some papers to be considered around these ideas. Any geneticist's interested in exploring these ideas further.
Charles R. , I make the most important points in my post, below, but I would be careful before trusting a tool like the one you have described completely.
This is because you are totally correct when you say this is a hot area in genetics. The thing is, it is SO hot that understanding of these regions has changed dramatically between the date of publication of the XYalign tool and today.
When I say "dramatically" I am not really exaggerating - in 2018-9 our human reference genome had about 30 Mbp for ChrY, but now we know that it is >60Mbp in length.
It's therefore important to say that, because our understanding was based on a reference genome that was in and of itself both inaccurate and incomplete, tools attempting to work with the regions of genome responsible for this inaccuracy benefit greatly from the newer information.
I am not sure I would trust data generated from NGS (short read, 2nd generation) alone in these very complex areas.
To explain this, I made a much longer post below.