18
J. Wakeley
genealogies. For example, Rauch and Bar-Yam (2004) studied the distribution of the
genetic “uniqueness” of a sample, defined as the length of the branch connecting that
sample to the rest of the gene genealogy. This distribution has an extremely long tail
because there is a chance the (n + 1)th sample will establish a new MRCA and thus
be markedly unique. Forward in time, the loss of such markedly unique lineages
causes neutral substitutions to accumulate in (non-Poisson) bursts (Watterson 1982;
Pfaffelhuber and Wakolbinger 2005). Greater statistical power to estimate θ may be
achieved by sampling more loci rather than increasing the sample size at a single
locus (Pluzhnikov and Donnelly 1996; Felsenstein 2006).
Predictions about site frequencies are obtained by considering mutations that
occur on branches in the coalescent tree that have i descendants in the sample. Fu
(1995) derived expressions for the expected values, variances, and covariances of
the site-frequency counts. Here, we will focus on the expected values, which are
E [ξ i ] =
θ
i
for i from 1 to n − 1. This simple relationship, for a sample of size n = 20, is graphed
as a “site-frequency spectrum” in Fig. 1.4a, which means that site frequencies are
plotted as expected proportions of all segregating sites. Figure 1.4b–d displays a
range of site-frequency spectra for three different models discussed in Sect. 1.5.
When depicted in this way, as expected proportions, the site-frequency spectrum
gives the probability of observing each type of polymorphism at a randomly chosen
SNP.
If the ancestral states at the sites of SNPs are not known, it is only possible to
discern the “folded” site-frequency patterns, which have expected values
E [η i ] = θ
1
i
+
1
n − i
1
δ i,n−i
for 1 ≤ i ≤ n/2. In the last term, δ i, n − i , is the Kronecker delta, so this term is a
correction for the case i = n − i, in which only one kind of SNP contributes to
η i . The full site-frequency spectrum is referred to as the “unfolded” site-frequency
spectrum.
None of these three measures of genetic variation depends on how variation
is arrayed along the sequences in the sample. All of them can be computed
by considering each SNP in isolation from all other SNPs. Patterns of linkage
disequilibrium between sites and the process of recombination that produces them
are the subject of Chap. 2. Although here our focus is on single loci without
recombination, it is important to note that all of the expected values given above hold
regardless of recombination. This is because the marginal coalescent process at each
site is the same as the corresponding single-locus coalescent process. However, the
variances given above hold only for single loci without recombination. Generally
speaking, recombination acts to decrease these variances because it introduces a
level of independence among sites.
J. Wakeley
genealogies. For example, Rauch and Bar-Yam (2004) studied the distribution of the
genetic “uniqueness” of a sample, defined as the length of the branch connecting that
sample to the rest of the gene genealogy. This distribution has an extremely long tail
because there is a chance the (n + 1)th sample will establish a new MRCA and thus
be markedly unique. Forward in time, the loss of such markedly unique lineages
causes neutral substitutions to accumulate in (non-Poisson) bursts (Watterson 1982;
Pfaffelhuber and Wakolbinger 2005). Greater statistical power to estimate θ may be
achieved by sampling more loci rather than increasing the sample size at a single
locus (Pluzhnikov and Donnelly 1996; Felsenstein 2006).
Predictions about site frequencies are obtained by considering mutations that
occur on branches in the coalescent tree that have i descendants in the sample. Fu
(1995) derived expressions for the expected values, variances, and covariances of
the site-frequency counts. Here, we will focus on the expected values, which are
E [ξ i ] =
θ
i
for i from 1 to n − 1. This simple relationship, for a sample of size n = 20, is graphed
as a “site-frequency spectrum” in Fig. 1.4a, which means that site frequencies are
plotted as expected proportions of all segregating sites. Figure 1.4b–d displays a
range of site-frequency spectra for three different models discussed in Sect. 1.5.
When depicted in this way, as expected proportions, the site-frequency spectrum
gives the probability of observing each type of polymorphism at a randomly chosen
SNP.
If the ancestral states at the sites of SNPs are not known, it is only possible to
discern the “folded” site-frequency patterns, which have expected values
E [η i ] = θ
1
i
+
1
n − i
1
δ i,n−i
for 1 ≤ i ≤ n/2. In the last term, δ i, n − i , is the Kronecker delta, so this term is a
correction for the case i = n − i, in which only one kind of SNP contributes to
η i . The full site-frequency spectrum is referred to as the “unfolded” site-frequency
spectrum.
None of these three measures of genetic variation depends on how variation
is arrayed along the sequences in the sample. All of them can be computed
by considering each SNP in isolation from all other SNPs. Patterns of linkage
disequilibrium between sites and the process of recombination that produces them
are the subject of Chap. 2. Although here our focus is on single loci without
recombination, it is important to note that all of the expected values given above hold
regardless of recombination. This is because the marginal coalescent process at each
site is the same as the corresponding single-locus coalescent process. However, the
variances given above hold only for single loci without recombination. Generally
speaking, recombination acts to decrease these variances because it introduces a
level of independence among sites.
