4 Types of Natural Selection and Tests of Selection
83
1989) (see above) can be used both to detect close to fixation or recently fixed
advantageous mutations or recently established balancing selection.
– The timing of selection. Different tests of positive selection aim at detecting either
hitchhiking or balancing selection at different times relative to the start and end
of the selective events. Ongoing selective sweeps not yet close to fixation or
balancing selection in its initial hitchhiking phase are best detected with tests
using haplotype structure in a single population. Recently completed selective
sweeps are best detected using either SFS-based tests or haplotype-based tests
using multiple populations (see above). Selection over extended time periods
involving multiple amino acid sites may best be detected using the McDonald–
Kreitman type test.
– The nature of the data. Tests of selection have strong requirements regarding
the data they need to be performed. Here we will discuss only the case where
genome-wide variation is available. The data used to test selection in the
genome-wide study usually consists either of sequencing data or of genome-wide
genotype calls obtained with a genotyping chip. Regardless of the kind of data
used, a fundamental parameter for selection tests is the number of individuals in
the dataset. Most tests need a minimum of 20 chromosomes or 10 individuals to
have the power to detect selection. The more individuals, the higher the power to
detect selection. Sequencing can be performed for each individual separately, or
for all individuals grouped together during the sequencing process. The latter is
called pooled sequencing.
Compared to sequencing, genome-wide genotyping is still much cheaper, but
there has recently been a clear movement toward adapting approaches to full
genome sequencing. Indeed, sequencing has become far more affordable. Importantly, genome-wide genotyping data suffers from the so-called ascertainment
bias (Clark et al. 2005). Ascertainment bias comes from the fact that only those
variants that were previously known are genotyped. This biases the genotyped
diversity in several ways. First, already known variants tend to be high-frequency
ones. This is a serious issue when using the SFS to detect selection, since the SFS
will be skewed toward intermediate frequencies. Such a skew is not trivial to
correct, since it varies along a chromosome with some loci being more affected
than others (Ramirez-Soriano and Nielsen 2009). Second, genotyped positions
are chosen so that they are regularly spaced along the genome. This regular
spacing removes the information on selection that could be extracted from the
local amount of variable sites (Clark et al. 2005). It also means that important
selected variants can be completely missed. A great amount of effort has been
done in the past 10 years to deal with ascertainment bias, but it is still a strong
limitation that explains why full sequencing is now favored over genotyping.
Sequencing is not free of local biases either. The main bias of sequencing
when trying to detect selection corresponds to the local variations in sequencing
depth (Abecasis et al. 2012). The lower the sequencing depth, the lower the
probability that all chromosomes have been sequenced in all individuals of the
studied samples. Incomplete sequencing affects the SFS by reducing the accuracy
of the estimated frequencies. At a given position the frequency of a variant in the
83
1989) (see above) can be used both to detect close to fixation or recently fixed
advantageous mutations or recently established balancing selection.
– The timing of selection. Different tests of positive selection aim at detecting either
hitchhiking or balancing selection at different times relative to the start and end
of the selective events. Ongoing selective sweeps not yet close to fixation or
balancing selection in its initial hitchhiking phase are best detected with tests
using haplotype structure in a single population. Recently completed selective
sweeps are best detected using either SFS-based tests or haplotype-based tests
using multiple populations (see above). Selection over extended time periods
involving multiple amino acid sites may best be detected using the McDonald–
Kreitman type test.
– The nature of the data. Tests of selection have strong requirements regarding
the data they need to be performed. Here we will discuss only the case where
genome-wide variation is available. The data used to test selection in the
genome-wide study usually consists either of sequencing data or of genome-wide
genotype calls obtained with a genotyping chip. Regardless of the kind of data
used, a fundamental parameter for selection tests is the number of individuals in
the dataset. Most tests need a minimum of 20 chromosomes or 10 individuals to
have the power to detect selection. The more individuals, the higher the power to
detect selection. Sequencing can be performed for each individual separately, or
for all individuals grouped together during the sequencing process. The latter is
called pooled sequencing.
Compared to sequencing, genome-wide genotyping is still much cheaper, but
there has recently been a clear movement toward adapting approaches to full
genome sequencing. Indeed, sequencing has become far more affordable. Importantly, genome-wide genotyping data suffers from the so-called ascertainment
bias (Clark et al. 2005). Ascertainment bias comes from the fact that only those
variants that were previously known are genotyped. This biases the genotyped
diversity in several ways. First, already known variants tend to be high-frequency
ones. This is a serious issue when using the SFS to detect selection, since the SFS
will be skewed toward intermediate frequencies. Such a skew is not trivial to
correct, since it varies along a chromosome with some loci being more affected
than others (Ramirez-Soriano and Nielsen 2009). Second, genotyped positions
are chosen so that they are regularly spaced along the genome. This regular
spacing removes the information on selection that could be extracted from the
local amount of variable sites (Clark et al. 2005). It also means that important
selected variants can be completely missed. A great amount of effort has been
done in the past 10 years to deal with ascertainment bias, but it is still a strong
limitation that explains why full sequencing is now favored over genotyping.
Sequencing is not free of local biases either. The main bias of sequencing
when trying to detect selection corresponds to the local variations in sequencing
depth (Abecasis et al. 2012). The lower the sequencing depth, the lower the
probability that all chromosomes have been sequenced in all individuals of the
studied samples. Incomplete sequencing affects the SFS by reducing the accuracy
of the estimated frequencies. At a given position the frequency of a variant in the
