1 Coalescent Models
21
others that test deviations from the unfolded site-frequency spectrum. Simonsen et
al. (1995) outlined a method of assessing the significance of such statistics which
accounts for the fact that P-values depend on an estimate of θ from the data. Fay
and Wu (2000) adapted one of Fu’s (1997) statistics as a test specifically for positive
selection. Their statistic, H, is sensitive to an excess of high-frequency SNPs, and
positive selection is one of a small number of deviations from the standard neutral
model that can cause such an excess. Achaz (2009) advanced the theory of devising
optimal statistics based on site frequencies, and Ferretti et al. (2010) extended
this approach to design optimal statistics for specific deviations from the standard
neutral model. Recently, Sainudiin and Véber (2018) described a way to compute
the likelihood of the full site-frequency spectrum of a sample at a locus without
recombination.
1.5
Extensions of the Standard Model
The strong simplifying assumptions of the standard neutral model—no selection,
constant population size over time, and no population structure—are appropriate
if the aim is to establish a null model to be tested. If one wishes to make more
detailed inferences about a range of biological phenomena, then coalescent models
must be extended to include those phenomena and their key parameters. Many
such extensions have been made, significantly broadening coalescent theory beyond
the standard neutral case. In this section, we will encounter examples of how two
important deviations from the Kingman coalescent, namely, changes in population
size over time and geographic population structure, can affect the site-frequency
spectrum and thus might be detected, for example, using Tajima’s D or related
statistics.
Figure 1.6 displays hypothetical gene genealogies for four different population
models, showing how the size and shape of gene genealogies depend on the details
of population structure and history. The standard neutral model (Fig. 1.6a), for
which expected site-frequency results are given in Fig. 1.4a, is compared to a
model of population growth (Fig. 1.6b), and two models of population structure:
divergence in isolation (Fig. 1.6c) and subdivision and migration (Fig. 1.6d). These
three types of deviations from the standard model are considered in detail below,
with expected site-frequency results given in the corresponding panels of Fig. 1.4.
A consideration of how natural selection affects site frequencies is taken up in
Chap. 4. Note that in some species, though probably not humans, extreme differences in offspring numbers among individuals can cause site-frequency patterns
that closely mimic those produced by natural selection. When the variance σ 2 of
the offspring-number distribution is very large, the Kingman coalescent may not
hold. Instead, gene genealogies may include multiple mergers of ancestral genetic
lineages (Möhle and Sagitov 2001). None of the predictions listed in Sect. 1.4 may
hold, and there may be a dramatic excess of high-frequency SNPs; e.g., see Sargsyan
and Wakeley (2008). Strong natural selection induces a very similar phenomena near
the locus at which selection acts (Durrett and Schweinsberg 2004; Etheridge et al.
21
others that test deviations from the unfolded site-frequency spectrum. Simonsen et
al. (1995) outlined a method of assessing the significance of such statistics which
accounts for the fact that P-values depend on an estimate of θ from the data. Fay
and Wu (2000) adapted one of Fu’s (1997) statistics as a test specifically for positive
selection. Their statistic, H, is sensitive to an excess of high-frequency SNPs, and
positive selection is one of a small number of deviations from the standard neutral
model that can cause such an excess. Achaz (2009) advanced the theory of devising
optimal statistics based on site frequencies, and Ferretti et al. (2010) extended
this approach to design optimal statistics for specific deviations from the standard
neutral model. Recently, Sainudiin and Véber (2018) described a way to compute
the likelihood of the full site-frequency spectrum of a sample at a locus without
recombination.
1.5
Extensions of the Standard Model
The strong simplifying assumptions of the standard neutral model—no selection,
constant population size over time, and no population structure—are appropriate
if the aim is to establish a null model to be tested. If one wishes to make more
detailed inferences about a range of biological phenomena, then coalescent models
must be extended to include those phenomena and their key parameters. Many
such extensions have been made, significantly broadening coalescent theory beyond
the standard neutral case. In this section, we will encounter examples of how two
important deviations from the Kingman coalescent, namely, changes in population
size over time and geographic population structure, can affect the site-frequency
spectrum and thus might be detected, for example, using Tajima’s D or related
statistics.
Figure 1.6 displays hypothetical gene genealogies for four different population
models, showing how the size and shape of gene genealogies depend on the details
of population structure and history. The standard neutral model (Fig. 1.6a), for
which expected site-frequency results are given in Fig. 1.4a, is compared to a
model of population growth (Fig. 1.6b), and two models of population structure:
divergence in isolation (Fig. 1.6c) and subdivision and migration (Fig. 1.6d). These
three types of deviations from the standard model are considered in detail below,
with expected site-frequency results given in the corresponding panels of Fig. 1.4.
A consideration of how natural selection affects site frequencies is taken up in
Chap. 4. Note that in some species, though probably not humans, extreme differences in offspring numbers among individuals can cause site-frequency patterns
that closely mimic those produced by natural selection. When the variance σ 2 of
the offspring-number distribution is very large, the Kingman coalescent may not
hold. Instead, gene genealogies may include multiple mergers of ancestral genetic
lineages (Möhle and Sagitov 2001). None of the predictions listed in Sect. 1.4 may
hold, and there may be a dramatic excess of high-frequency SNPs; e.g., see Sargsyan
and Wakeley (2008). Strong natural selection induces a very similar phenomena near
the locus at which selection acts (Durrett and Schweinsberg 2004; Etheridge et al.
