3 Analysis of Population Structure
55
2002, for a review on some of these methods). The dissimilarity measure of the
individuals in the original data is the scientist’s choice, and some examples of
useful dissimilarity measures were presented in the previous section. The number of
dimensions in the plotting space is chosen in advance and is usually low, to facilitate
observation and interpretation of the data.
3.2.3 Ancestry Component Estimation with Few Model
Assumptions
The program STRUCTURE (Pritchard et al. 2000) implements a method that
makes very few assumptions about the data, and it was one of the first tools that
utilized the power of multiple markers for inference. This approach and the many
ensuing approaches have become standard in population structure investigations
and population-genetic studies in general. STRUCTURE-like methods infer a
predefined number of ancestry components (K) among individuals, based on
genotype frequencies. Each individual’s genotype is assigned to one of K number
of clusters with a certain probability. In the first implementation, STRUCTURE
searched for the assignment of individuals that minimizes deviation from HardyWeinberg equilibrium (see Box 3.1) and linkage equilibrium in each of the K
clusters and allowed individuals to be admixed and to have membership proportions
to more than one of the K clusters (Pritchard et al. 2000). Population structure is
then visible in the dataset as individuals that are closely related having a greater
proportion of their genome assigned to the same cluster/s than individuals that
are not. This approach analyzed single markers separately and then added up the
information to produce a global estimate for each individual. Information about
the relative positions of markers to each other was not used and was considered
to be segregating independently. This approach works well for low-density marker
sets but is less suitable for the high density and full genome datasets that are
available today. In the 2003 update of the STRUCTURE algorithm (Falush et al.
2003), sites/markers are not required to be independent, and correlations between
subsequent markers due to admixture events are explicitly modeled. This allowed
for individual ancestry estimates, known as local ancestry estimates, where the
ancestry of chromosomal chunks can be traced along the chromosomes. It also
introduced a simplistic model (the F-model, originally described in Nicholson et
al. (2002)) to account for correlations of allele frequencies between populations.
Although a clearly unrealistic model, it improved the performance of the algorithm
considerably (Falush et al. 2003).
Box 3.1 Hardy–Weinberg Equilibrium (HWE)
An assumption of random mating is that the probability to produce viable
offspring is equal for all possible pairs of individuals drawn from the
(continued)
55
2002, for a review on some of these methods). The dissimilarity measure of the
individuals in the original data is the scientist’s choice, and some examples of
useful dissimilarity measures were presented in the previous section. The number of
dimensions in the plotting space is chosen in advance and is usually low, to facilitate
observation and interpretation of the data.
3.2.3 Ancestry Component Estimation with Few Model
Assumptions
The program STRUCTURE (Pritchard et al. 2000) implements a method that
makes very few assumptions about the data, and it was one of the first tools that
utilized the power of multiple markers for inference. This approach and the many
ensuing approaches have become standard in population structure investigations
and population-genetic studies in general. STRUCTURE-like methods infer a
predefined number of ancestry components (K) among individuals, based on
genotype frequencies. Each individual’s genotype is assigned to one of K number
of clusters with a certain probability. In the first implementation, STRUCTURE
searched for the assignment of individuals that minimizes deviation from HardyWeinberg equilibrium (see Box 3.1) and linkage equilibrium in each of the K
clusters and allowed individuals to be admixed and to have membership proportions
to more than one of the K clusters (Pritchard et al. 2000). Population structure is
then visible in the dataset as individuals that are closely related having a greater
proportion of their genome assigned to the same cluster/s than individuals that
are not. This approach analyzed single markers separately and then added up the
information to produce a global estimate for each individual. Information about
the relative positions of markers to each other was not used and was considered
to be segregating independently. This approach works well for low-density marker
sets but is less suitable for the high density and full genome datasets that are
available today. In the 2003 update of the STRUCTURE algorithm (Falush et al.
2003), sites/markers are not required to be independent, and correlations between
subsequent markers due to admixture events are explicitly modeled. This allowed
for individual ancestry estimates, known as local ancestry estimates, where the
ancestry of chromosomal chunks can be traced along the chromosomes. It also
introduced a simplistic model (the F-model, originally described in Nicholson et
al. (2002)) to account for correlations of allele frequencies between populations.
Although a clearly unrealistic model, it improved the performance of the algorithm
considerably (Falush et al. 2003).
Box 3.1 Hardy–Weinberg Equilibrium (HWE)
An assumption of random mating is that the probability to produce viable
offspring is equal for all possible pairs of individuals drawn from the
(continued)
