3 Analysis of Population Structure
57
A subclass of population admixture models can utilize the spatial coordinates
of sampled individuals (BAPS (Corander et al. 2003), TESS (Chen et al. 2007),
GENELAND (Guillot et al. 2005)). While in STRUCTURE-like methods, the
assignment to a population is independent and identical among all individuals in
the dataset, this class of methods takes into account (a priori) the spatial distribution
of individuals and aims to detect genetic discontinuities in space. Since geographic
spatial correlation is often present among individuals and populations, it can be
useful to incorporate spatial coordinates into the population structure analysis.
Recent attempts to incorporate geographic information into population structure
estimations involve approaches that use Wishart distributions (a generalization of
multidimensional gamma distributions) to model genetic similarity as a function of
spatial distance. In uniform isolation-by-distance scenarios, genetic distances visualized in two dimensions should mirror the samples/individuals in geographic space.
Migration and admixture and hinders to gene flow would disturb this correlation.
In one approach, implemented in the software SpaceMix, a covariance model of
genetic data is used to build maps of the geographic positions of the populations,
but distances are distorted according to inferred rates of gene flow (Bradburd et
al. 2016). Barriers to gene flow result in larger distances between groups, while
migration and admixture can be identified as abnormal strong covariances over long
distances. The inferred admixture is then estimated and represented as “arrows,” on
a generated map, from the source population to the recipient population. Another
approach based on the Wishart distribution, EEMS (Petkova et al. 2016), uses
pairwise genetic similarities among populations and estimates a surface map of
effective migration rates. The effective migration rates are scaled by effective
population sizes under an equilibrium model. These methods that incorporate spatial
information may highlight important features of population structure that might
have remained undetected using other, spatially “blind,” methods for inferring
population structure. Two other methods that use F ST measures between populations
to identify violations of isolation-by-distance patterns have also been developed
(Duforet-Frebourg and Blum 2014; Jay et al. 2013).
For our example dataset, we ran 10 iterations in the program ADMIXTURE
at K = 2 to K = 5. The iterations at each value of K were then compared to
detect different clustering solutions using the program CLUMPP (Jakobsson and
Rosenberg 2007). For K = 2 to K = 4, all 10 iterations arrived at very similar
solutions, and the combined output is shown in Fig. 3.4 (visualized with the program
DISTRUCT (Rosenberg 2004)). The analysis shows clustering of sub-Saharan
Africans (orange component) and clustering of Europeans and North Africans (blue
component) at K = 2, although Yoruba individuals show a small fraction of ancestry
from the blue component, and the Mozabites show a small ancestry fraction from the
orange component. At K = 3, the Yoruba obtains its own cluster (green), while the
Mozabite still clusters with the French but showing some ancestry fraction from the
green component. The Mozabite forms its own group at K = 4, but some individuals
showed shared ancestry with Yoruba and French. For K = 5, there was no common
solution, and each of the 10 iterations had a different solution (this pattern can be
seen clearly from the similarity matrix output of CLUMPP). The lack of a “common
Précédent

- 63/236

Suivant