26
J. Wakeley
We can understand the pattern in Fig. 1.4d by imagining a mixture of patterns like
the one in Fig. 1.4c, depending on the specific outcome of migration and coalescence
in the recent ancestry of the sample. For example, if one ancestral lineage migrates
out of the sampled deme and the other 19 ancestral lineages coalesce within it, then
the gene genealogy will resemble one for which two demes were sampled with
n 1 = 1 and n 2 = 19. This would cause mutants counts ξ 1 and ξ 19 to be inflated. Fig.
1.6d shows an analogous scenario for a sample of size six in a five-deme model, in
which the counts ξ 2 and ξ 4 would be inflated.
Figure 1.5 illustrates that such deviations can be detected using Tajima’s D.
For example, if we adopt the lower 2.5% critical value of −1.803 and the upper
97.5% critical value of 2.001 for n = 20 from Table 2 in Tajima (1989), the one
million pseudo-datasets from the standard neutral coalescent that yielded Fig. 1.5a
reject the null model 2.0% of the time in the negative direction and 1.1% in the
positive direction. In contrast, the pseudo-datasets that yielded Fig. 1.5b, which were
generated under the same kind of many-demes model that gave the site-frequency
spectrum in Fig. 1.4d (but with E[π w ] = 10 to allow comparison with Fig. 1.5a)
reject the null model 11.4% of the time in the negative direction and 18.8% in the
positive direction.
1.6
Conclusion: Current Challenges of Big Data
Now is a particularly exciting time in human population genetics. The field is awash
in data, with major efforts such as the Simons Genome Diversity Project (Mallick
et al. 2016), the 1000 Genomes Project (The 1000 Genomes Project Consortium
2015), and the UK Biobank (Bycro et al. 2018) promising that, soon, many millions
of genomes will be available for study. The continued relevance of the models
presented here may be seen in the recent papers by Kelleher et al. (2019) and Speidel
et al. (2019). These present new methods for the population-genetic analysis of very
large numbers of genomes. With the caveat that at the genomic scale, it is crucial
to include recombination (see Chap. 2), parts of the analyses in both Kelleher et al.
(2019) and Speidel et al. (2019) rely on the standard neutral coalescent model. The
aim in both works is to infer the ordered series of mutation events and coalescent
events at loci across the human genome (recall the importance-sampling methods
described in Sect. 1.4.2). In doing so, both works use the techniques of Li and
Stephens (2003), which extend the importance-sampling method of Stephens and
Donnelly (2000) to account for recombination. The results of Kelleher et al. (2019)
and Speidel et al. (2019) provide first-pass estimates of ancient relationships and
associated mutations among humans across the genome (Harris 2019).
The aim of this chapter has been to describe the foundational models of
coalescent theory. They are simplified models which capture the effects of neutral
mutation, reproduction, and genetic transmission in shaping distributions of genetic
diversity. The simplest model, the standard neutral coalescent, assumes a single
well-mixed population of constant size, but extensions to include changes in
population size over time and idealized kinds of population structure were also
J. Wakeley
We can understand the pattern in Fig. 1.4d by imagining a mixture of patterns like
the one in Fig. 1.4c, depending on the specific outcome of migration and coalescence
in the recent ancestry of the sample. For example, if one ancestral lineage migrates
out of the sampled deme and the other 19 ancestral lineages coalesce within it, then
the gene genealogy will resemble one for which two demes were sampled with
n 1 = 1 and n 2 = 19. This would cause mutants counts ξ 1 and ξ 19 to be inflated. Fig.
1.6d shows an analogous scenario for a sample of size six in a five-deme model, in
which the counts ξ 2 and ξ 4 would be inflated.
Figure 1.5 illustrates that such deviations can be detected using Tajima’s D.
For example, if we adopt the lower 2.5% critical value of −1.803 and the upper
97.5% critical value of 2.001 for n = 20 from Table 2 in Tajima (1989), the one
million pseudo-datasets from the standard neutral coalescent that yielded Fig. 1.5a
reject the null model 2.0% of the time in the negative direction and 1.1% in the
positive direction. In contrast, the pseudo-datasets that yielded Fig. 1.5b, which were
generated under the same kind of many-demes model that gave the site-frequency
spectrum in Fig. 1.4d (but with E[π w ] = 10 to allow comparison with Fig. 1.5a)
reject the null model 11.4% of the time in the negative direction and 18.8% in the
positive direction.
1.6
Conclusion: Current Challenges of Big Data
Now is a particularly exciting time in human population genetics. The field is awash
in data, with major efforts such as the Simons Genome Diversity Project (Mallick
et al. 2016), the 1000 Genomes Project (The 1000 Genomes Project Consortium
2015), and the UK Biobank (Bycro et al. 2018) promising that, soon, many millions
of genomes will be available for study. The continued relevance of the models
presented here may be seen in the recent papers by Kelleher et al. (2019) and Speidel
et al. (2019). These present new methods for the population-genetic analysis of very
large numbers of genomes. With the caveat that at the genomic scale, it is crucial
to include recombination (see Chap. 2), parts of the analyses in both Kelleher et al.
(2019) and Speidel et al. (2019) rely on the standard neutral coalescent model. The
aim in both works is to infer the ordered series of mutation events and coalescent
events at loci across the human genome (recall the importance-sampling methods
described in Sect. 1.4.2). In doing so, both works use the techniques of Li and
Stephens (2003), which extend the importance-sampling method of Stephens and
Donnelly (2000) to account for recombination. The results of Kelleher et al. (2019)
and Speidel et al. (2019) provide first-pass estimates of ancient relationships and
associated mutations among humans across the genome (Harris 2019).
The aim of this chapter has been to describe the foundational models of
coalescent theory. They are simplified models which capture the effects of neutral
mutation, reproduction, and genetic transmission in shaping distributions of genetic
diversity. The simplest model, the standard neutral coalescent, assumes a single
well-mixed population of constant size, but extensions to include changes in
population size over time and idealized kinds of population structure were also
