48
P. Sjödin et al.
and equals for all possible pairs of individuals drawn from the population (often,
depending on the species, given that a pair consists of a female and a male).
Random mating in this sense is, however, rarely an accurate representation of
reality since most populations have a spatial distribution that affects the mating
probability of a random pair of individuals. Hence, if by “structured population,” we
refer to a population that is not randomly mating, then essentially all populations are
structured, and a dual categorization of “structured populations” vs “unstructured
populations” is not very useful. Instead, it may be better to think of any population
as structured and attempt to quantify the degree of structure. The level of population
structure can also (often) be accounted for in downstream analyses. In practice, there
are often groups of individuals for which no structure is detectable (with the data
and methods at hand), and these groups can be regarded as unstructured for most
intents and purposes.
It should be noted that this way of thinking about populations does not reflect
common practice in which the definition of populations is typically subjective,
based on, for example, linguistic, cultural, ethnic, and/or the geographic location
of sampled individuals. Moreover, almost all population structure analyses are
necessarily based on samples, and detection of stratification within the sample does
not necessarily reflect biological populations. For instance (following Pritchard
et al. (2000)), imagine a species that lives on a continuous plane but has a low
dispersal rate, so allele frequencies vary continuously across the plane. A few
clustered sampling points will result in a signal of (a few) clustered genetic groups,
which does not give an accurate description of the biological reality. This example
illustrates that studies of populations are always indirect and limited to information
contained in samples and sensitive to sampling biases.
Accounting for population structure is often crucial in order to reduce both type
I and type II errors in statistical analyses of genetic data. For instance, it has been
shown that not accounting for population structure can result in spurious signals
in association mapping studies and will thus invalidate standard tests (Ewens and
Spielman 1995; Pritchard et al. 2000). It is also important to account for population
structure in forensic applications like DNA fingerprinting to estimate the probability
of random individuals matching a particular profile (Balding and Nichols 1994,
1995; Foreman et al. 1997; Roeder et al. 1998).
Since the first historical opportunity of quantifying molecular genetic variation
until today’s (almost) complete genome sequencing, numerous molecular techniques have been used to genotype individuals, which in turn can be used to
investigate population structure. Early strategies involved typing the human blood
groups (Landsteiner and Weiner 1940), followed by the development of allozymes
(Lewontin and Hubby 1966), and various forms of DNA fragment length assays,
including microsatellites that are abundant in eukaryote genomes (Katti et al. 2001).
The sequencing of the human genome (Venter et al. 2001) led to access to a large
number of human single-nucleotide polymorphisms (SNPs) as well as the human
genome sequence. Various studies have employed novel molecular techniques
to investigate human population structure, including classical markers (CavalliSforza et al. 1994), mitochondrial genome (Cann et al. 1987), microsatellites
Précédent

- 54/236

Suivant