24
J. Wakeley
structure, and these labels affect the rates of coalescence, then the lineages are
not exchangeable, and modeling gene genealogies become more complicated. In
the case of population subdivision, either with or without migration, the chance
of coalescence is greater for pairs of lineages in the same subpopulation than for
pairs of lineages in different subpopulations. This can only be modeled by explicitly
keeping track of the locations of ancestral lineages as they are followed backward
in time.
Other complications arise because structured populations may contain many
subpopulations, which may be of different sizes and between which any number
of complex patterns of migration might exist. A general model of D subpopulations,
or “demes” as they are often called, would have D 2 parameters: D deme sizes and
D(D − 1) migration rates. In addition, it is not clear what sort of simplified limiting
models should be developed for structured populations. Some populations might
comprise a small number of very large demes, while others might comprise a large
number of small demes. It could be that the very idea of demes/subpopulations is
inapplicable, rather than that the population is continuously distributed across its
habitat.
Accordingly, a number of different coalescent models of geographic structure
have been developed—these are reviewed in Hein et al. (2005) and Wakeley
(2009)—and the choice of model must depend on the species being studied. Wright
(1931) introduced the island model of population subdivision with migration, which
has been the source of a great number of other models and methods of data
analysis. Herbots (1997) and Notohara (1990) described the general mathematical
coalescent approach to these discrete-deme models, following Takahata (1988). In
these models, the deme sizes are assumed to be large, like the population size in the
Kingman coalescent (N → ∞). The migration rates are assumed to be small. They
are treated in the same manner that mutation is treated in the limit leading to the
Kingman coalescent.
The resulting structured coalescent model allows the straightforward derivation
of useful expressions concerning genetic variation. For instance, consider a simple
version of the island model in which all D demes are of the same (haploid) size
N, migration between all pairs of demes occurs with the same per-generation
probability m, and reproduction occurs by haploid Wright–Fisher sampling. Then, if
π w and π b are the average number of differences between pairs of sequences from
the same deme (i.e., “within”) and the average number of differences between pairs
of sequences from two different demes (i.e., “between”), it can be shown that
E [π w ] = θD
and
E [π b ] = θD
1 +
1
M
J. Wakeley
structure, and these labels affect the rates of coalescence, then the lineages are
not exchangeable, and modeling gene genealogies become more complicated. In
the case of population subdivision, either with or without migration, the chance
of coalescence is greater for pairs of lineages in the same subpopulation than for
pairs of lineages in different subpopulations. This can only be modeled by explicitly
keeping track of the locations of ancestral lineages as they are followed backward
in time.
Other complications arise because structured populations may contain many
subpopulations, which may be of different sizes and between which any number
of complex patterns of migration might exist. A general model of D subpopulations,
or “demes” as they are often called, would have D 2 parameters: D deme sizes and
D(D − 1) migration rates. In addition, it is not clear what sort of simplified limiting
models should be developed for structured populations. Some populations might
comprise a small number of very large demes, while others might comprise a large
number of small demes. It could be that the very idea of demes/subpopulations is
inapplicable, rather than that the population is continuously distributed across its
habitat.
Accordingly, a number of different coalescent models of geographic structure
have been developed—these are reviewed in Hein et al. (2005) and Wakeley
(2009)—and the choice of model must depend on the species being studied. Wright
(1931) introduced the island model of population subdivision with migration, which
has been the source of a great number of other models and methods of data
analysis. Herbots (1997) and Notohara (1990) described the general mathematical
coalescent approach to these discrete-deme models, following Takahata (1988). In
these models, the deme sizes are assumed to be large, like the population size in the
Kingman coalescent (N → ∞). The migration rates are assumed to be small. They
are treated in the same manner that mutation is treated in the limit leading to the
Kingman coalescent.
The resulting structured coalescent model allows the straightforward derivation
of useful expressions concerning genetic variation. For instance, consider a simple
version of the island model in which all D demes are of the same (haploid) size
N, migration between all pairs of demes occurs with the same per-generation
probability m, and reproduction occurs by haploid Wright–Fisher sampling. Then, if
π w and π b are the average number of differences between pairs of sequences from
the same deme (i.e., “within”) and the average number of differences between pairs
of sequences from two different demes (i.e., “between”), it can be shown that
E [π w ] = θD
and
E [π b ] = θD
1 +
1
M
