16
J. Wakeley
but the expressions are cumbersome. Interested readers should consult the review
by Tavaré (1984) or the textbooks by Hein et al. (2005) and Wakeley (2009).
1.4.2 Levels and Patterns of Genetic Variation
The introduction of the coalescent was revolutionary in population genetics because
it provided a way to compute the probability of a dataset. This probability,
also known as the likelihood, is the foundation of rigorous statistical inference.
Population-genetic data are complicated, with nested patterns of variation among
subsets of the sample as in Fig. 1.1, and no general analytical results are available
for the likelihood. Inference has proceeded by the numerical computation of
likelihoods, using simulations to sample gene genealogies from the coalescent prior
distribution. Even this is quite complicated due to the enormity of the sample space
of gene genealogies.
Two ways of accounting for the unknown gene genealogy were developed in
the 1990s: importance sampling and Markov chain Monte Carlo (MCMC). If we
knew the order of mutation events and coalescent events, as in Fig. 1.1, but not
the times, we could compute the likelihood by multiplying the probabilities of the
ordered events. For example, under the infinite-sites model, all likelihoods include
the familiar probability of identity by descent, 1/(θ + 1), originally due to Malécot
(1946), because for all gene genealogies, the final event is that two ancestral lineages
coalesce at the MRCA before either of them mutates. Importance-sampling methods
average these products of probabilities over possible orderings of events (Griffiths
and Tavaré 1994, 1996; Stephens and Donnelly 2000; Wu 2010). Alternatively, if
we knew the tree structure and the coalescence times, we could model the process
of mutation along the branches of the tree in computing the likelihood. MCMC
methods do this and average over the underlying trees and times (Kuhner et al.
1995; Kuhner 2006; Beerli 2006; Hey and Nielsen 2004, 2007; Drummond et al.
2012).
There has been a growing trend to make inferences based on “summary
statistics” rather than the full data, often within the framework of approximate
Bayesian computation (Beaumont 2010; Alvarado-Serrano and Hickerson 2016).
The summary-statistic approach to inference reduces the dimensionality of the data,
ideally to a small set of simpler measures of genetic variation which are highly
informative with regard to a set of parameters or phenomena of interest. As with
importance sampling and MCMC, summary-statistic approaches use coalescent
models to average over gene genealogies. Coalescent theory is also used to make
predictions about summary statistics, a number of which (e.g., heterozygosity) have
also been important historically in population genetics.
Three kinds of summary statistics have been well studied in a variety of settings.
These are segregating sites, average pairwise differences, and site frequencies,
which are defined as follows for a sample from a single population. The number
of segregating sites, S, is the number of polymorphisms, e.g., SNPs in a dataset
of DNA sequences. The average number of pairwise differences, which is denoted
Précédent

- 23/236

Suivant