78
D. Enard
initially introduced it in order to detect loci with an excess of rare deleterious
mutations. However, because of the popularity of genome scans for positive
selection, it has been used more widely to detect selective sweeps. The principle of
the Tajima’s D test is to contrast two well-known estimators of genetic diversity,
namely, π and S on the folded SFS to detect a departure from neutrality. π is
the average of differences between each pair of sequences in the sample studied.
S is the number of segregating variants. When a selective sweep is very near or
right after fixation, variation at perfectly linked sites right next to the selected
mutation has been wiped out. Further from the selected mutation, recombination
maintains variation, but many preexisting older intermediate-frequency alleles have
been swept away from the population. At the same time, there is a constant input of
new, low-frequency alleles. This results in an excess of young, rare alleles compared
to the neutral SFS. The strength of the excess of rare variants reflects the strength
and speed of the selective episode. The missing intermediate-frequency alleles do
not affect S strongly as they represented only a small proportion of all segregating
variants in the first place. Unlike S, π is strongly affected by the lack of intermediatefrequency alleles. As a reminder, π is the average of differences between each pair
of sequences in the sample studied. Intermediate-frequency alleles in the folded
SFS result in differences between many pairs of sequences in the sample, whereas
rare alleles result in differences between very few pairs of sequences. π is therefore
far more sensitive to intermediate-frequency variants than S is. Tajima’s D is the
difference between π and an estimate of the mutation rate based on S, normalized
by the standard deviation of the difference. Under perfectly neutral conditions,
the difference is expected to be null. When a sweep is right before or right after
fixation, Tajima’s D is negative: π is strongly reduced because of the missing
intermediate frequency alleles, while S is less affected. It is important to note that
any demographic expansion or bottleneck that induces a local excess of rare alleles
can create spurious detections of sweeps when using Tajima’s D. Tajima’s D is
also sensitive to the excess of rare alleles induced by background selection due to
weakly deleterious mutations. Tajima’s D can also be used to detect recent balancing
selection, when there is an excess of intermediate-frequency alleles across a wide
region surrounding the balanced allele not yet broken down by recombination. In
this case, Tajima’s D is strongly positive.
4.2.1.2 Fay and Wu’s H
Tajima’s D inspired a number of other statistics similarly based on the SFS (Achaz
2009). In an attempt to create a statistic more sensitive to positive selection at
the exclusion of other confounding processes, Fay and Wu created the statistic H
(Fay and Wu 2000), usually called Fay and Wu’s H. H uses the unfolded SFS,
which means an out-group has to be used to polarize alleles. The rationale is
to detect selective sweeps using intermediate- and high-frequency alleles of the
unfolded SFS, at the exclusion of rare alleles. Indeed, the main effect of population
expansions, bottlenecks, and background selection on the SFS is to create an excess
of young rare alleles. By focusing on intermediate- and high-frequency variants,
H is expected to be more robust to these confounding factors. In brief, H is the
D. Enard
initially introduced it in order to detect loci with an excess of rare deleterious
mutations. However, because of the popularity of genome scans for positive
selection, it has been used more widely to detect selective sweeps. The principle of
the Tajima’s D test is to contrast two well-known estimators of genetic diversity,
namely, π and S on the folded SFS to detect a departure from neutrality. π is
the average of differences between each pair of sequences in the sample studied.
S is the number of segregating variants. When a selective sweep is very near or
right after fixation, variation at perfectly linked sites right next to the selected
mutation has been wiped out. Further from the selected mutation, recombination
maintains variation, but many preexisting older intermediate-frequency alleles have
been swept away from the population. At the same time, there is a constant input of
new, low-frequency alleles. This results in an excess of young, rare alleles compared
to the neutral SFS. The strength of the excess of rare variants reflects the strength
and speed of the selective episode. The missing intermediate-frequency alleles do
not affect S strongly as they represented only a small proportion of all segregating
variants in the first place. Unlike S, π is strongly affected by the lack of intermediatefrequency alleles. As a reminder, π is the average of differences between each pair
of sequences in the sample studied. Intermediate-frequency alleles in the folded
SFS result in differences between many pairs of sequences in the sample, whereas
rare alleles result in differences between very few pairs of sequences. π is therefore
far more sensitive to intermediate-frequency variants than S is. Tajima’s D is the
difference between π and an estimate of the mutation rate based on S, normalized
by the standard deviation of the difference. Under perfectly neutral conditions,
the difference is expected to be null. When a sweep is right before or right after
fixation, Tajima’s D is negative: π is strongly reduced because of the missing
intermediate frequency alleles, while S is less affected. It is important to note that
any demographic expansion or bottleneck that induces a local excess of rare alleles
can create spurious detections of sweeps when using Tajima’s D. Tajima’s D is
also sensitive to the excess of rare alleles induced by background selection due to
weakly deleterious mutations. Tajima’s D can also be used to detect recent balancing
selection, when there is an excess of intermediate-frequency alleles across a wide
region surrounding the balanced allele not yet broken down by recombination. In
this case, Tajima’s D is strongly positive.
4.2.1.2 Fay and Wu’s H
Tajima’s D inspired a number of other statistics similarly based on the SFS (Achaz
2009). In an attempt to create a statistic more sensitive to positive selection at
the exclusion of other confounding processes, Fay and Wu created the statistic H
(Fay and Wu 2000), usually called Fay and Wu’s H. H uses the unfolded SFS,
which means an out-group has to be used to polarize alleles. The rationale is
to detect selective sweeps using intermediate- and high-frequency alleles of the
unfolded SFS, at the exclusion of rare alleles. Indeed, the main effect of population
expansions, bottlenecks, and background selection on the SFS is to create an excess
of young rare alleles. By focusing on intermediate- and high-frequency variants,
H is expected to be more robust to these confounding factors. In brief, H is the
