52
P. Sjödin et al.
we created an IBS distance matrix in Plink and used the R-package ape to construct
an NJ tree. We see that the individuals cluster quite well according to their sample
locations, except for the Mozabites that contain the French sample as a subset.
3.2.2 Principal Component Analysis and Related Approaches
A single individual can be represented by a position in a multidimensional space
where each locus characterizes one dimension. The number of dimensions is
therefore very large if we have access to information from many sites; oftentimes,
however, many of these dimensions are correlated. A number of methods based on
linear algebra aim at finding the best way to summarize and visualize the data in
a reduced number of dimensions that capture the greatest axes of variation. These
methods, which are typically agnostic to the underlying model of genetic variation,
can potentially reveal inherent population structure in a set of sampled individuals.
We will describe one of the most widely used method for initial data exploration—
principal component analysis (PCA)—and then give a brief overview of a few other
related approaches.
The principle of PCA is straightforward: finding and ordering orthogonal axes
(or principal components, PCs) that capture the variation of the sample so that the
first PC represents the axis of greatest variation in the data, the second PC represents
the axis of greatest remaining variation when the data is projected orthogonally to
the first PC, and so on, down to the last PC where all variation has been taken into
account. Consider, for example, n SNP-loci. Each individual is represented by a
vector of n values, with 0 if homozygous for the reference allele, 1 if homozygous
for the alternative allele, and 0.5 if heterozygous. PCA performs a rotation of the
original n-dimensional orthogonal base, where each of the n loci represents one
dimension, into a new orthogonal base, formed by linear combinations of the loci,
and defining directions called principal components. The first PC is the direction
that maximally explains the variance among individuals when projected into a 1dimensional space. Together with the first PC, the second PC defines the plane
that maximizes the variance of the individuals when projected into a 2-dimensional
orthogonal space, and more generally, together with the first k-1 PCs, PC k defines
an orthogonal space that maximizes the variance of the individuals when projected
into a k-dimensional space. Each PC explains a proportion of the total variance, with
the first PC explaining the most variance and the last PC explaining the least.
Applied to our example data of four populations, PCA reveals differences
between the groups (Fig. 3.2). We see that the first and second PCs explain 10.8%
and 5.5% of the total sample variance, respectively. The first PC separates the subSaharan African populations from the European and the North African population;
the second PC separates the Southern African San and the West African Yoruba,
with the French and the Mozabite in-between. The first two PCs together show
the Mozabite individuals in-between the European population and the West African
population, suggesting historical models for the ancestry of the Mozabite such as
(1) a history of admixture between European and West African groups, (2) shared
P. Sjödin et al.
we created an IBS distance matrix in Plink and used the R-package ape to construct
an NJ tree. We see that the individuals cluster quite well according to their sample
locations, except for the Mozabites that contain the French sample as a subset.
3.2.2 Principal Component Analysis and Related Approaches
A single individual can be represented by a position in a multidimensional space
where each locus characterizes one dimension. The number of dimensions is
therefore very large if we have access to information from many sites; oftentimes,
however, many of these dimensions are correlated. A number of methods based on
linear algebra aim at finding the best way to summarize and visualize the data in
a reduced number of dimensions that capture the greatest axes of variation. These
methods, which are typically agnostic to the underlying model of genetic variation,
can potentially reveal inherent population structure in a set of sampled individuals.
We will describe one of the most widely used method for initial data exploration—
principal component analysis (PCA)—and then give a brief overview of a few other
related approaches.
The principle of PCA is straightforward: finding and ordering orthogonal axes
(or principal components, PCs) that capture the variation of the sample so that the
first PC represents the axis of greatest variation in the data, the second PC represents
the axis of greatest remaining variation when the data is projected orthogonally to
the first PC, and so on, down to the last PC where all variation has been taken into
account. Consider, for example, n SNP-loci. Each individual is represented by a
vector of n values, with 0 if homozygous for the reference allele, 1 if homozygous
for the alternative allele, and 0.5 if heterozygous. PCA performs a rotation of the
original n-dimensional orthogonal base, where each of the n loci represents one
dimension, into a new orthogonal base, formed by linear combinations of the loci,
and defining directions called principal components. The first PC is the direction
that maximally explains the variance among individuals when projected into a 1dimensional space. Together with the first PC, the second PC defines the plane
that maximizes the variance of the individuals when projected into a 2-dimensional
orthogonal space, and more generally, together with the first k-1 PCs, PC k defines
an orthogonal space that maximizes the variance of the individuals when projected
into a k-dimensional space. Each PC explains a proportion of the total variance, with
the first PC explaining the most variance and the last PC explaining the least.
Applied to our example data of four populations, PCA reveals differences
between the groups (Fig. 3.2). We see that the first and second PCs explain 10.8%
and 5.5% of the total sample variance, respectively. The first PC separates the subSaharan African populations from the European and the North African population;
the second PC separates the Southern African San and the West African Yoruba,
with the French and the Mozabite in-between. The first two PCs together show
the Mozabite individuals in-between the European population and the West African
population, suggesting historical models for the ancestry of the Mozabite such as
(1) a history of admixture between European and West African groups, (2) shared
