4 Types of Natural Selection and Tests of Selection
79
difference between π and θ H where π is sensitive to intermediate-frequency alleles
as explained in the description of Tajima’s D, and θ H is a measure of diversity that is
sensitive to high-frequency alleles. Right before or after fixation, many intermediate
frequency alleles have been fixed, and π is low. At the same time, young, derived
alleles linked to the selected mutation were rapidly driven to high frequencies.
There is an excess of young, high-frequency derived alleles, and θ H is high. The
difference H = π – θ H is therefore strongly negative soon before and after the
fixation of a selected mutation. Because the effect of population expansions is to
increase the supply of new, rare alleles, population expansions have no effect on
H. Conversely, population bottlenecks can affect H. The rapid rise of a specific
haplotype to a high frequency can occur by chance during a bottleneck. This mimics
a selective sweep and creates both an excess of high-frequency derived alleles and
a paucity of intermediate alleles on a local scale. Fay and Wu’s H is robust to
background selection driven by strongly deleterious mutations (Zeng et al. 2006).
The robustness of the statistic to background selection driven by weakly deleterious
mutations remains to be evaluated.
4.2.1.3 Likelihood Ratio Tests
In order to achieve greater statistical power to detect selective sweeps, several
authors have developed Likelihood Ratio Tests (LRTs) based on the SFS. The
advantage of such tests is to consider the entire SFS rather than just using parts
of it as for Tajima’s D or Fay and Wu’s H. Further, LRTs compare the likelihoods
of the observed SFS under a selective sweep model and a neutral model, with the
likelihood of the observed SFS under a neutral model. As it provides an explicit test
of selection, it should be more powerful and robust than simple tests of neutrality.
Under both sweep and neutral models, it is possible to estimate the probability of
the observed frequency for each allele of the tested locus. The likelihood of the SFS
under each model is then calculated by multiplying all the individual probabilities
for all the sites present in the tested locus. Once the likelihoods are known both
with and without sweeps, the likelihood ratio can be used to decide whether or
not the neutral hypothesis can be rejected. There are two distinct approaches to
estimate the likelihood of the SFS under the neutral hypothesis. The first approach
consists of using a classical neutral model under panmictic assumptions (Kim and
Stephan 2002; Li and Stephan 2005). A major issue with this neutral model is
that it does not take demographic perturbations into account. Consequently, this
approach is very sensitive to demographic events such as population expansions or
bottlenecks (Jensen et al. 2005). The second approach, called CLR for composite
likelihood ratio (Nielsen et al. 2005), is to estimate the likelihood of the SFS at a
local scale based on the global, genome-wide SFS as the neutral model. Compared
to the previous classical neutral model, this empirical neutral model is a good
approximation of the average expected SFS given the past demographic history
of the population studied. Although this approach does integrate the systematic,
average influence of demography in the neutral model, it still does not account
for the increased variance in the local SFS along chromosomes expected after
Précédent

- 85/236

Suivant