3 Analysis of Population Structure
49
(Rosenberg et al. 2002), SNPs (Jakobsson et al. 2008; Li et al. 2008), and complete
genomes (1000 Genomes, Mallick et al. 2016). Although these different types of
data have different properties, the general study of population structure follows
straightforward principles of genetic variation, and the difference in datatype can
be accounted for by using different assumptions on the mutation model (see e.g.,
Veeramah and Hammer 2014).
We will review some methods to quantify structure within a population. We will
(primarily) assume biallelic markers, where only two states exist that give rise to
three possible genotypes. These methods are often part of an initial exploratory
step in order to get a better picture of how and to what extent the sample and
the population have been influenced by population structure. First, we will present
methods where all individuals are treated equally and no a priori information is used
for clustering individuals or parts of individual genomes. We will then discuss some
methods to contrast clusters of individuals and how this can be used in more explicit
demographic modeling. To illustrate concepts and methods, we will analyze a subset
of the HGDP data, which has been thoroughly investigated in previous studies (Cann
et al. 2002; Li et al. 2008; Jakobsson et al. 2008). This example dataset consists of
individuals from Africa—the West African Yoruba, the Southern African San, and
the North African Mozabite—and from Europe (France).
3.2
Individual-Based and Unsupervised Methods for Inferring
Population Structure
3.2.1 Tree Construction Methods at the Individual Level
Genetic distance is a traditional measure of differentiation in population genetics.
Distances are calculated between each pair of individuals (can also be calculated
between groups of individuals; see below) and are represented by a pairwise
distance matrix. Distance matrices can be visualized by various approaches such
as the multidimensional statistics discussed in the next section or in the form of
graphs/trees.
A common distance measure between a pair of individuals is the identity by
state (IBS) measure. IBS examines biallelic SNPs between two individuals and puts
them into one of three categories: identical = 1 (e.g., for the genotypes AA and
AA, BB and BB, and AB and AB, where A and B denote the two alleles), oneallele-shared = 0.5 (i.e., AA and AB; AB and BB), and no-allele-shared = 0 (i.e.,
AA and BB). The state-values are then averaged over all loci to provide genomewide pairwise IBS similarity values between 1 and 0 for all individual pairwise
comparisons. This is summarized in an individual similarity matrix of which 1-IBS
will give the distance matrix.
Distance measures based on substitution models for DNA and protein sequence
evolution have also been developed. The evolutionary distance between a pair of
sequences is measured by the number of nucleotide (or amino acid) substitutions
occurring between them. The p-distance is the simplest model and is based on
the proportion (p) of nucleotide sites at which two sequences being compared
49
(Rosenberg et al. 2002), SNPs (Jakobsson et al. 2008; Li et al. 2008), and complete
genomes (1000 Genomes, Mallick et al. 2016). Although these different types of
data have different properties, the general study of population structure follows
straightforward principles of genetic variation, and the difference in datatype can
be accounted for by using different assumptions on the mutation model (see e.g.,
Veeramah and Hammer 2014).
We will review some methods to quantify structure within a population. We will
(primarily) assume biallelic markers, where only two states exist that give rise to
three possible genotypes. These methods are often part of an initial exploratory
step in order to get a better picture of how and to what extent the sample and
the population have been influenced by population structure. First, we will present
methods where all individuals are treated equally and no a priori information is used
for clustering individuals or parts of individual genomes. We will then discuss some
methods to contrast clusters of individuals and how this can be used in more explicit
demographic modeling. To illustrate concepts and methods, we will analyze a subset
of the HGDP data, which has been thoroughly investigated in previous studies (Cann
et al. 2002; Li et al. 2008; Jakobsson et al. 2008). This example dataset consists of
individuals from Africa—the West African Yoruba, the Southern African San, and
the North African Mozabite—and from Europe (France).
3.2
Individual-Based and Unsupervised Methods for Inferring
Population Structure
3.2.1 Tree Construction Methods at the Individual Level
Genetic distance is a traditional measure of differentiation in population genetics.
Distances are calculated between each pair of individuals (can also be calculated
between groups of individuals; see below) and are represented by a pairwise
distance matrix. Distance matrices can be visualized by various approaches such
as the multidimensional statistics discussed in the next section or in the form of
graphs/trees.
A common distance measure between a pair of individuals is the identity by
state (IBS) measure. IBS examines biallelic SNPs between two individuals and puts
them into one of three categories: identical = 1 (e.g., for the genotypes AA and
AA, BB and BB, and AB and AB, where A and B denote the two alleles), oneallele-shared = 0.5 (i.e., AA and AB; AB and BB), and no-allele-shared = 0 (i.e.,
AA and BB). The state-values are then averaged over all loci to provide genomewide pairwise IBS similarity values between 1 and 0 for all individual pairwise
comparisons. This is summarized in an individual similarity matrix of which 1-IBS
will give the distance matrix.
Distance measures based on substitution models for DNA and protein sequence
evolution have also been developed. The evolutionary distance between a pair of
sequences is measured by the number of nucleotide (or amino acid) substitutions
occurring between them. The p-distance is the simplest model and is based on
the proportion (p) of nucleotide sites at which two sequences being compared
