1 Coalescent Models
9
Fortunately, it appears that inferences based on the standard coalescent model
may often involve little error because the process of coalescence on a fixed pedigree is practically indistinguishable from the standard neutral coalescent process
(Wakeley et al. 2012). This result comes from simulations of large well-mixed
populations, so it is important to note that it may not hold for all possible extensions
of coalescent modeling. For example, in those finely divided, structured populations,
often referred to as “meta-populations” (Hanski and Gaggiotti 2004), the sizes of
local subpopulations may be small, and averaging over pedigrees could give highly
inaccurate results.
Standard coalescent models become accurate for large well-mixed populations
because the ancestries of all present-day individuals overlap broadly (Chang 1999;
Rohde et al. 2003). If all ancestors are distinct, then every individual will have
2 g pedigree ancestors in generation g in the past. Thus, just 40 generations, or
perhaps 1000 years ago, we should each have more than one trillion ancestors.
However, according to Fig. 1 of Keinan and Clark (2012), the number of people
alive 1000 years ago was only about 100 million. For each of us, our >10 12 expected
pedigree ancestors must all map onto 10 8 actual pedigree ancestors. This causes
a huge degree of overlap of our ancestries. For a Wright–Fisher population of
large constant size N, Chang (1999) found that by 1.77log 2 N generations ago the
population is divided neatly into two groups: a fraction (~0.7698) who are ancestors
of every present-day individual and a fraction (~0.2302) who have no descendants
today. For perspective, 1.77log 2 N is roughly 35 generations for a population of size
N = 10 6 and 47 generations for N = 10 8 .
Roughly speaking, it is because of this broad overlap of pedigree ancestries in
the relatively recent past that coalescent models based on the incorrect assumption
of homogeneity of coalescent probabilities over time actually make reasonable
predictions about the distribution of gene genealogies within fixed population pedigrees, at least for large well-mixed populations. Of course, the distribution of gene
genealogies within fixed population pedigrees is not identical to the distribution
of gene genealogies under the standard neutral coalescent. But the differences are
primarily restricted to the past log 2 N generations, at which point there is a rapid
transition to the type of homogeneous, essentially pedigree-independent behavior
found in the standard model (Wakeley et al. 2012).
Random samples from large well-mixed populations are very unlikely to include
closely related individuals, so the typical effect of the population pedigree is to
bar coalescence in the very recent past. But since log 2 N generations is much less
than the timescale of the coalescent process, i.e., N generations, the effects of
population pedigrees on gene genealogies will often be negligible. This provides
some justification for the common practice of discarding individuals when high
levels of relatedness are detected in population-genetic data (Rosenberg 2006).
However, it may be preferable to account for recent pedigrees explicitly, particularly
in structured populations (Wilton et al. 2017). In cases where the pedigree itself is
of interest, Ko and Nielsen (2019) describe how it can be estimated from genetic
data.
Précédent

- 16/236

Suivant