1 Coalescent Models
25
where the parameter M = 2Nm is the scaled migration rate. This simple model can
be extended to include other modes of reproduction or diploidy, as in the Kingman
coalescent, by replacing N with N e in both θ and M.
The expression for E[π w ] says that the expected level of genetic variation within
demes is identical to the expected level in a single population of the same total size,
ND. The same is not true of the variance (not shown), which depends inversely
on M. The expression for E[π b ] says that the expected level of genetic variation
between demes is increased by an amount inversely proportional to the migration
parameter M. When M is small, the population is expected to contain high levels
of variation, and the difference between π b and π w is expected to be great. This
is the basis of F ST as a measure of the degree of population subdivision (Slatkin
1991). It is important to note that the scaled migration rate M may be large, so that
little evidence of subdivision is apparent, even if the per-generation probability of
migration is small.
These results had been known previously (Li 1976; Slatkin 1987; Strobeck 1987),
but the introduction of the structured coalescent greatly facilitated the development
of sophisticated methods of inference for structured populations, akin to those
mentioned in Sect. 1.4.3, where the model is employed to average over gene
genealogies in the computation of the likelihood (de Iorio et al. 2005; Beerli 2006).
Similar methods have been proposed for cases of nonequilibrium migration, in
which two or more populations descend from a single ancestral population (Hey
and Nielsen 2004, 2007; Wilkinson-Herbots 2008; Hey 2010).
Population subdivision can have a dramatic effect on site frequencies. Figure
1.4C shows the site-frequency spectrum for a sample of size n = 20 for a
hypothetical case of hidden population structure. Specifically, the sample contains
n 1 = 6 sequences from one population and n 2 = 14 from the other population under
the isolation model of Takahata and Nei (1985) in which two populations split from
a common ancestral population at some time in the past and after that exchanged
no migrants. The same mutation rate θ = 1 was used for all three populations, and
the split time was assumed to be t = 1, measured on the coalescent timescale. In
this case, there is a tendency for the gene genealogy to be composed of two subtrees with 6 and 14 tips connected by a long internal branch, with mutations on
this branch contributing to ξ 6 and ξ 14 . Figure 1.6c depicts a similar scenario for a
sample of total size six.
Figure 1.4d shows the site-frequency spectrum for a sample of size n = 20 taken
from a single deme in the island model with many demes and a migration rate of
M = 1. Here, in the recent past, ancestral genetic lineages not only coalesce within
the deme from which they were sampled but also migrate to other demes. When
all remaining ancestral lineages are in separate demes, the process of coalescence
is dependent on migration events that bring ancestral lineages together into the
same deme, where they have a chance to coalesce. Both the branching pattern and
the lengths of coalescent intervals differ from those of standard coalescent gene
genealogies. For M = 1, this results in the slightly U-shaped site-frequency spectrum
shown in Fig. 1.4d.
25
where the parameter M = 2Nm is the scaled migration rate. This simple model can
be extended to include other modes of reproduction or diploidy, as in the Kingman
coalescent, by replacing N with N e in both θ and M.
The expression for E[π w ] says that the expected level of genetic variation within
demes is identical to the expected level in a single population of the same total size,
ND. The same is not true of the variance (not shown), which depends inversely
on M. The expression for E[π b ] says that the expected level of genetic variation
between demes is increased by an amount inversely proportional to the migration
parameter M. When M is small, the population is expected to contain high levels
of variation, and the difference between π b and π w is expected to be great. This
is the basis of F ST as a measure of the degree of population subdivision (Slatkin
1991). It is important to note that the scaled migration rate M may be large, so that
little evidence of subdivision is apparent, even if the per-generation probability of
migration is small.
These results had been known previously (Li 1976; Slatkin 1987; Strobeck 1987),
but the introduction of the structured coalescent greatly facilitated the development
of sophisticated methods of inference for structured populations, akin to those
mentioned in Sect. 1.4.3, where the model is employed to average over gene
genealogies in the computation of the likelihood (de Iorio et al. 2005; Beerli 2006).
Similar methods have been proposed for cases of nonequilibrium migration, in
which two or more populations descend from a single ancestral population (Hey
and Nielsen 2004, 2007; Wilkinson-Herbots 2008; Hey 2010).
Population subdivision can have a dramatic effect on site frequencies. Figure
1.4C shows the site-frequency spectrum for a sample of size n = 20 for a
hypothetical case of hidden population structure. Specifically, the sample contains
n 1 = 6 sequences from one population and n 2 = 14 from the other population under
the isolation model of Takahata and Nei (1985) in which two populations split from
a common ancestral population at some time in the past and after that exchanged
no migrants. The same mutation rate θ = 1 was used for all three populations, and
the split time was assumed to be t = 1, measured on the coalescent timescale. In
this case, there is a tendency for the gene genealogy to be composed of two subtrees with 6 and 14 tips connected by a long internal branch, with mutations on
this branch contributing to ξ 6 and ξ 14 . Figure 1.6c depicts a similar scenario for a
sample of total size six.
Figure 1.4d shows the site-frequency spectrum for a sample of size n = 20 taken
from a single deme in the island model with many demes and a migration rate of
M = 1. Here, in the recent past, ancestral genetic lineages not only coalesce within
the deme from which they were sampled but also migrate to other demes. When
all remaining ancestral lineages are in separate demes, the process of coalescence
is dependent on migration events that bring ancestral lineages together into the
same deme, where they have a chance to coalesce. Both the branching pattern and
the lengths of coalescent intervals differ from those of standard coalescent gene
genealogies. For M = 1, this results in the slightly U-shaped site-frequency spectrum
shown in Fig. 1.4d.
