3 Analysis of Population Structure
65
simulated data are compared to the values of these summary statistics observed
from the real data. The simulations are performed by drawing model parameters
from prior distributions and then choosing those simulations that best mimic the
real data. The distribution of the model parameters in this chosen set of best fitting
simulations can then be used to estimate the model parameters. Different models can
then be contrasted using Bayes factors. The ABC approach has proven very flexible
for inferring model parameters, and there is an active community developing novel
and faster algorithms (e.g., Pudlo et al. 2016; Csilléry et al. 2012).
3.4
Summary and Guidelines
Good practice when investigating a population-genetic dataset for population structure is to start by visualizing the data in a way that reveals the inherent characteristics
of the data. To get an indication of whether or not the data contains different
groups; PCA, simple tree-building methods, and “STRUCTURE-like” approaches
are all good tools for initial data exploration. Once some overview of the data is
obtained, we can start building simple models to investigate additional hierarchical
patterns and characteristics of the underlying demographic history of the sample
and population/s. More detailed hypotheses can subsequently be scrutinized using
more explicit and advanced models. The more accurate models of demography we
can infer, the better we understand the underlying processes shaping the population
genetic patterns of variation. Ultimately, this understanding may allow in-depth
analysis of the genetic architecture of traits and patterns of selection impacting the
genome (Li et al. 2012), after controlling for patterns caused by demographic history
manifesting as population structure.
References
Alexander DH, Novembre J, Lange K (2009) Fast model-based estimation of ancestry in unrelated
individuals. Genome Res 19:1655–1664
Balding DJ, Nichols RA (1994) DNA profile match probability calculation: how to allow for
population stratification, relatedness, database selection and single bands. Forensic Sci Int
64:125–140
Balding DJ, Nichols RA (1995) A method for quantifying differentiation between populations at
multi-allelic loci and its implications for investigating identity and paternity. Genetica 96:3–12
Beaumont MA, Zhang W, Balding DJ (2002) Approximate Bayesian computation in population
genetics. Genetics 162:2025–2035
Becquet C, Przeworski M (2007) A new approach to estimate parameters of speciation models
with application to apes. Genome Res 17:1505–1519
Bhatia G, Patterson N, Sankararaman S, Price AL (2013) Estimating and interpreting FST: the
impact of rare variants. Genome Res 23:1514–1521
Bradburd GS, Ralph PL, Coop GM (2016) A spatial framework for understanding population
structure and admixture. PLoS Genet 12:e1005703
Cann RL, Stoneking M, Wilson AC (1987) Mitochondrial DNA and human evolution. Nature
325:31–36
Précédent

- 71/236

Suivant