1 Coalescent Models
11
When the Kingman coalescent is the limiting ancestral process, it is useful to
refer to the coalescent effective population size N e (Sjödin et al. 2005), which is
given by N/σ 2 in Kingman’s derivation from the general Cannings’ model, by the
familiar 2N in the diploid Wright–Fisher model, under both the monoecious model
and the dioecious model with equal numbers of males and females, and by 1/c N in
general.
Note that the statement N → ∞ does not refer to changes in the size of the
population. The population size N is assumed to be constant over time in the
standard neutral coalescent model (later, one may relax this assumption). The limit
simply means that we consider a series of such (constant-size) populations, with the
aim of identifying the dominant behavior of the ancestral process when N is very
large.
The standard neutral coalescent has been shown to be robust to many deviations
from Kingman’s initial assumptions (Möhle 1998a). It applies when generations
are overlapping and to populations of diploid, biparental organisms. The latter case
requires mathematical formalism beyond what Kingman used. This was developed
in a pair of papers by Möhle which treated partial selfing (Möhle 1998b) and
diploid, biparental inheritance (Möhle 1998c). In all these cases, the derivation of
the coalescent begins with the description of an expected, single-generation process,
which is the average over all possible outcomes of reproduction or over the pedigree.
1.3.1 The Sampling Structure of Coalescent Gene Genealogies
The end product of these calculations is a continuous-time model of the ancestral
genetic process which begins with the n genetic lineages of the present-day sample
and proceeds back into the past. Each pair of lineages coalesces independently with
a rate equal to one, so that the total rate is i(i − 1)/2 when there are i ancestral
lineages. Again, i(i − 1)/2 is the total number of pairs of lineages that can coalesce.
Thus, the total rate of coalescence is higher when there are more lineages available
to coalesce with each other. Coalescent events occur between randomly chosen pairs
of lineages at randomly (exponentially) distributed times. The process is stopped
when the last two ancestral lineages coalesce into a single lineage, the MRCA of
the sample.
One run of this process produces a random-joining tree with associated branch
lengths determined by the series of exponentially distributed coalescence times,
which is taken to represent a single gene genealogy sampled from the distribution
of all possible gene genealogies under the model. Multiple independent runs are
used to represent collections of gene genealogies at multiple unlinked or effectively
unlinked loci. Gene genealogies vary quite dramatically, in both branching structure
and coalescence times (reflected in the heights of the genealogies). This is shown
in Fig. 1.3 which displays ten randomly generated gene genealogies for a sample of
size n = 20 under the standard neutral coalescent. A key purpose of coalescent
theory is to model variation in gene genealogies, as in Fig. 1.3, reflecting the
randomness of the evolutionary process.
Précédent

- 18/236

Suivant