Hi dthorbur,
Thanks so much for your reply! The purpose of the phylogeny is a little convoluted, but it is not for a study with a deep evolutionary focus. Essentially, the population with GBS data is a USDA gene bank collection that is readily and globally accessible to researchers, while the WGS data represents a collection that is not accessible outside it's host nation. I wish to find a set of USDA accessions that represent the major clades found within the WGS accessions. We intend to construct a pangenome resource based on these target accessions.
Your idea of selecting a group of loci with similar depth distributions sounds good, that is definitely something I will try. What order of magnitude would you suggest a 'handful of loci' should be? A hundred per chromosome? A thousand? Unfortunately my study organism doesn't have established sets of loci used for population analyses, so I don't think I can follow-up on that angle.
If you have any further advice I'd really appreciate it!
Thanks again, Max