1 Coalescent Models
13
modeled quite accurately as a continuous-time Markov process or sometimes simply
as a Poisson counting process. A four-state Markov process is appropriate for a
mutation in DNA, with its four nucleotides. When a Poisson counting process is
used as an approximation, it is also often assumed that at most one mutation can
occur per site. This is known as the infinite-sites (or infinitely-many-sites) mutation
model.
The mutation rate for a genetic locus is typically denoted θ /2, the mutation
parameter θ being proportional to the product of the population size, N, and neutral
mutation rate per generation, u. In general, θ = 2N e u, with θ = 4Nu in the diploid
Wright–Fisher model. Technically, θ is assumed to exist in the limit N → ∞, but
less formally, the model is valid when N is large and u is small. With θ defined
this way, the number of mutations on a branch or branches of total length t follows
a Poisson distribution with expected value tθ /2. The critical feature of the infinitesites model is that each mutation creates a unique polymorphic site. Thus, for a
given nucleotide site in the genome, at most, one mutation can have occurred in the
history of the sample. This chapter will focus exclusively on this mutation model,
which is due in this form (i.e., without recombination) to Watterson (1975). The
infinite-sites model is a reasonable starting approximation for human autosomal
genetic diversity, because only about 1/1000 nucleotide sites are polymorphic when
two human genomes are compared (Cargill et al. 1999; Stephens et al. 2001) and
only about 1/500 SNPs show more than two bases segregating (Hodgkinson and
Eyre-Walker 2010).
A large number of four-state models have been put forward to represent DNA
mutations, the HKY85 model being one of the most commonly used (Hasegawa et
al. 1985). Models of “stepwise” mutation have also been added to the coalescent
in order to account for variation in repeat sequences, such as microsatellite loci
(Valdes et al. 1993). In general, mutation is a time-inhomogeneous process and
must be modeled separately along each branch, forward in time starting from the
MRCA or root of the tree. A number of simpler, “parent-independent” mutation
models have been employed as approximations; for example, see Stephens and
Donnelly (2003) and Fearnhead (2006). The infinite-sites model considered here
and the infinite-alleles model used by Ewens (1972) are special cases of parentindependent mutation.
1.4
Fundamental Predictions for Single Loci in Well-Mixed
Populations
The mathematical convenience of the standard neutral coalescent, the ease with
which it may be applied, and all of the detailed predictions one can make using
it follow from three key properties of standard neutral gene genealogies:
• The branching structure of a coalescent tree is determined by randomly joining
pairs of ancestral lineages until the MRCA is reached
Précédent

- 20/236

Suivant