8 Inferring Human Demographic History from Genetic Data
193
with greater hitchhiking effects on the X chromosome. This result is also apparent
from linkage-disequilibrium-based estimates of X and autosome population sizes
(Lohmueller et al. 2010). In addition, systematic differences between continental
groups are likely the result of shared demographic processes such as a population
bottleneck associated with the exodus of modern humans out of Africa (Arbiza et
al. 2014). In summary, while it is tempting to speculate on the relative prevalence
of polygyny vs. polyandry across human history, it is extremely difficult to separate
out the effects of other population processes that differentially affect autosomal and
sex-linked levels of diversity.
8.6
Estimating Demographic Parameters
Recently, researchers have started to develop computational and statistical methods
for jointly estimating multiple demographic parameters (e.g., split times and
migration rates) from DNA sequence data. Generally, these methods must navigate
a tradeoff between statistical rigor and biological realism, since the statistically
“optimal” approach of full maximum likelihood on autosomal sequence data from
multiple individuals is computationally infeasible for the foreseeable future. Below,
some of the basic approaches researchers have used are outlined, as well as some
recent applications to human data.
The major computational burden of inference methods comes from the modeling
of intragenic recombination. One way of reducing the computational burden is
to assume a model with no recombination (e.g., Nielsen and Wakeley 2001;
Drummond and Rambaut 2007; Gronau et al. 2011) and to apply this model
to mtDNA data, data from short autosomal regions without visible evidence of
recombination, or single diploid genome sequences. While the no-recombination
assumption reduces the usefulness of these methods, this is partially counteracted by
the ability to employ full-likelihood techniques (generally using Bayesian Markov
chain Monte Carlo algorithms) to efficiently utilize the information contained in
the data. For example, Gignoux and colleagues used a large mtDNA sequence data
set to infer the rates of recent population growth in different human populations
(Gignoux et al. 2011). Their results suggest that most of the population growth
happened within the past 8000 years, consistent with recent analyses of autosomal
sequence data (described above in Sect. 8.3).
Researchers have also tried the opposite tactic of assuming free recombination
between all sites (Nielsen 2000; Marth et al. 2004; Garrigan 2009; Gutenkunst et
al. 2009; Nielsen et al. 2009; Kamm et al. 2017, 2019). Under this assumption,
sequence data can be summarized by the site frequency spectrum (SFS, the distribution of the number of SNPs with different allele frequencies), and the expected
relative values for the SFS can be calculated computationally or analytically for
complex demographic models (e.g., Garrigan 2009; Gutenkunst et al. 2009; Bhaskar
et al. 2015; Kamm et al. 2017, 2019). These methods have the benefit of being able
to handle genome-wide polymorphism data in a computationally efficient manner
but at the cost of making a biologically unrealistic assumption and ignoring an
Précédent

- 195/236

Suivant