4
J. Wakeley
Winther et al. (2015) discuss three common uses of “population”: in mathematical models, in the laboratory, and in the wild. The populations in this chapter are
of the first kind. They are theoretical constructs to be applied for the sake of better
understanding. Their application may lead either to further advances in modeling or
to new hypotheses about populations of actual organisms including ourselves. The
framework is statistical and involves sampling from populations. It should be borne
in mind that “the sample” in what follows means genetic data taken from a number
of individuals.
As a matter of perspective, it is important to recognize the surprising overall
truth about human genetic variation—that there is very little of it compared to what
is found in most species (Leffler et al. 2012). Further, the degree of substructure
among humans is remarkably low (Rosenberg et al. 2005). As a first approximation,
it is not uncommon or unreasonable to compare global patterns of human genetic
variation to the predictions for a single, well-mixed population.
1.2
Introduction: Gene Genealogies Within a Population
or Species
Population-genetic datasets, with their typically complex and interesting patterns
of polymorphism among the DNA sequences, haplotypes or genotypes in the
sample, are the result of an equally complex and interesting set of ancestral genetic
processes. Each single-nucleotide polymorphism (SNP) reflects the specific patterns
of descent from the ancestors of the sample and the mutation(s) at that nucleotide
site during genetic transmission. Patterns of descent from ancestors are influenced
by the random processes of genetic transmission and a host of demographic
processes which may include natural selection, population growth, and population
structure. The data consist only of patterns of polymorphism, and the challenge is
to use these to make whatever inferences we can about the underlying processes.
Forgetting mutations for the moment, the term gene genealogy refers to the
pattern of genetic ancestry among the members of a sample at a single-nucleotide
site or a genetic locus made up of a non-recombining sequence of sites. If there is
intra-locus recombination, then gene genealogies at different sites may be different
(see Chap. 2). In this chapter, gene genealogies are considered without intra-locus
recombination. Under mild restrictions on the sample size, the population size, and
the demography of the species, gene genealogies may be depicted accurately as
rooted, bifurcating trees, with the samples at the tips and the most recent common
ancestor, or MRCA, of the sample at the root. The branches represent the genetic
lineages ancestral to the sample.
Figure 1.1 shows a hypothetical dataset and a corresponding gene genealogy. For
a real dataset, the gene genealogy would be unknown, but it is clear from Fig. 1.1
that the structure of the gene genealogy is a very strong determinant of the patterns
of mutations (e.g., the frequencies) in the sample. Fig. 1.1 depicts the simple case in
which every mutation in the ancestry of the sample occurs at a different nucleotide
Précédent

- 11/236

Suivant