12
J. Wakeley
Fig. 1.3 Ten independently generated gene genealogies for a sample of size n = 20, produced
using a Mathematica Demonstrations Project “Coalescent Gene Genealogies” written by John
Hawks
The standard neutral coalescent provides a prior distribution of gene genealogies
which can be invoked (logically before a sample is taken) to make predictions about
expected patterns of genetic variation or for purposes of statistical inference from
data. For example, Huff et al. (2010) used a simple result from coalescent theory,
due to Tajima (1983), to identify loci in a pair of human genomes that had twice the
average coalescence time of randomly chosen loci, and employed these older loci to
make inferences about ancient human effective population sizes. General methods
of statistical inference for larger samples, such as those mentioned in Sect. 1.4.3,
treat gene genealogies as missing data and average over them using the coalescent
prior.
1.3.2 Including Mutations in the Coalescent
The lineages of the gene genealogy represent all the opportunity for mutations
in the ancestry of the sample: any polymorphisms in the data must be the result
of mutations that occurred along the branches of the gene-genealogical tree.
Predictions about genetic variation and inferences from genetic data cannot be made
unless mutations are included in the model. Fortunately, this is straightforward in
the standard neutral coalescent. By definition, neutral genetic variation does not
affect the probabilities of reproduction or the distribution of the number of offspring
per individual, so mutation and coalescence can be treated separately. In particular,
conditional on the gene-genealogical tree, mutations occur independently along
each branch.
Because the timescale of coalescence is in units of N generations, each branch
in the tree represents a huge number of opportunities for a mutation to occur. Then,
because the probability of mutation per generation is very small, mutation may be
Précédent

- 19/236

Suivant