Warning: I'm at home without access to Nature and I don't have the methods in front of me, so I can only answer based on what is in your question.
Are you sure that you accept this result? Just because something is published in Nature, it doesn't mean that it is true. You need to look at everything with a very critical eye. I don't know the authors or the paper, though they have good reputations and I'm sure that they are careful scientists, but in complex data analysis that involves a lot of somewhat arbitrary choices, there's a lot of opportunity for confirmation bias to sneak in.
Here's a few of things to think about when interpreting Figure 2 from Novembre et al:
-
The distribution of simulated phenotypes was (linearly) scaled, rotated and flipped to make it correspond as closely as possible to the map of Europe. Even so, there are areas where this correspondence breaks down, for example, the ES/PT distributions are almost completely overlapping. Also, why are only ~1,400 individuals shown in the figure, not the full 3,000 individuals mentioned in the abstract? How was this subset selected?
-
How were genetic distance between two individuals was calculated? Is it simply the covariance of allelic frequencies across all 500k SNPs, or is there some pre-selection of SNPs and/or some sort of unusual distance measure? If it is the latter, before you accept Figure 2, you need to convince yourself that there was no feedback between the result and the choice of distance measure.
-
Where was the phenotyping done? Was it all in the same lab, or were there different labs depending on country of origin of the sample? Batch effects can depend on geographically distributed quantities like humidity.
-
The last thing to look closely at how the "country of origin" was originally assigned to each sample. Was there some sort of pre-selection for exemplar samples?
In terms of PCA, if genetic distance does scale with geographic distance, and this is the major source of variation among the samples, then (as the authors point out) it is not at all surprising that the first two principal components are linear functions of latitude and longitude. This might even work if there the relationship between geographic and genetic distance was non-linear but monotonically increasing, so long as the genetic distances were not very large: most manifolds are locally linear.
Interesting topic, but not really bioinformatics, IMHO. This pattern is in the data, and not caused by the analysis methods.
I think it is an interesting and thoughtfully posed question. Importantly, it indirectly addresses one of the key analysis issues for any genetic association problem: population sub-structure confounding with endpoints. I think this is something that bioinformatics analysts should be aware of if they're going to work with GWAS data.