80
D. Enard
bottlenecks, and false positives due to demography are still an issue (Huber et al.
2014).
4.2.2 Tests of Selection Based on Haplotypes
4.2.2.1 Extended Haplotype Homozygosity (EHH)
The EHH (extended haplotype homozygosity) test was introduced in 2002 by
Sabeti et al. (2002) to test the hypothesis of positive selection of two variants
protecting against malaria at the G6PD and CD40L loci. Both alleles decrease
the risk of malaria by approximately 50% in heterozygous individuals. In African
regions where malaria is endemic, G6PD-202A is at a frequency of 20% despite
the fact that it causes mild G6PD deficiency. This clearly suggests that increased
resistance to malaria has given G6PD-202A heterozygous individuals a strong
selective advantage despite the adverse effects of the allele. If G6PD-202A has
driven a selective sweep, it is a partial ongoing selective sweep. Statistics aimed at
detecting complete or near-complete sweeps such as Tajima’s D or Fay and Wu’s H
have no power in such a case. If selection for G6PD-202A has been strong enough,
the allele should be associated with an unusually long haplotype block that has not
been broken down by recombination. The EHH statistic was designed to detect such
long haplotypes. It is built by first defining core haplotypes. Core haplotypes are
combinations of nearby alleles that are in perfect linkage (i.e., no recombination
has occurred between the alleles). For all possible pairs of chromosomes with the
same core haplotype, EHH is calculated as the proportion of pairs with perfect
homozygosity over a physical distance x from the core haplotype. The distance x
is fixed and largely defined by the size of the locus that was sequenced around the
candidate allele. A relative EHH is then calculated as the ratio of EHH for the core
haplotype of interest to the EHH for all other core haplotypes grouped together. The
relative EHH is meant to normalize for recombination. Note however that there is
always more power to detect selection in regions with low recombination rates and
a greater risk of false positives (O’Reilly et al. 2008; Ferrer-Admetlla et al. 2014).
The observed relative EHH can finally be compared with a distribution of simulated
neutral relative EHH. Sabeti et al. (2002) successfully used the EHH test to detect
strong partial sweeps at the G6PD and CD40L loci. To this date, the G6PD locus is
one of the best examples of an ongoing, partial sweep in the human genome.
4.2.2.2 Integrated Haplotype Score (iHS)
A major limitation of the EHH test is that it uses an arbitrary distance to the core
haplotype. One can imagine the curve of EHH as a function of the distance to the
core haplotype. The area under the curve captures more information about haplotype
structure and represents a less arbitrary statistic than EHH measured for a fixed
distance to the core haplotype. In this respect, the integrated Haplotype Score (iHS)
(Voight et al. 2006) is an improved, integrated, and standardized version of EHH.
Unlike the classic EHH, iHS is oriented based on the derived or ancestral status
of alleles. Single ancestral or derived alleles are used as core haplotypes, and iHS
D. Enard
bottlenecks, and false positives due to demography are still an issue (Huber et al.
2014).
4.2.2 Tests of Selection Based on Haplotypes
4.2.2.1 Extended Haplotype Homozygosity (EHH)
The EHH (extended haplotype homozygosity) test was introduced in 2002 by
Sabeti et al. (2002) to test the hypothesis of positive selection of two variants
protecting against malaria at the G6PD and CD40L loci. Both alleles decrease
the risk of malaria by approximately 50% in heterozygous individuals. In African
regions where malaria is endemic, G6PD-202A is at a frequency of 20% despite
the fact that it causes mild G6PD deficiency. This clearly suggests that increased
resistance to malaria has given G6PD-202A heterozygous individuals a strong
selective advantage despite the adverse effects of the allele. If G6PD-202A has
driven a selective sweep, it is a partial ongoing selective sweep. Statistics aimed at
detecting complete or near-complete sweeps such as Tajima’s D or Fay and Wu’s H
have no power in such a case. If selection for G6PD-202A has been strong enough,
the allele should be associated with an unusually long haplotype block that has not
been broken down by recombination. The EHH statistic was designed to detect such
long haplotypes. It is built by first defining core haplotypes. Core haplotypes are
combinations of nearby alleles that are in perfect linkage (i.e., no recombination
has occurred between the alleles). For all possible pairs of chromosomes with the
same core haplotype, EHH is calculated as the proportion of pairs with perfect
homozygosity over a physical distance x from the core haplotype. The distance x
is fixed and largely defined by the size of the locus that was sequenced around the
candidate allele. A relative EHH is then calculated as the ratio of EHH for the core
haplotype of interest to the EHH for all other core haplotypes grouped together. The
relative EHH is meant to normalize for recombination. Note however that there is
always more power to detect selection in regions with low recombination rates and
a greater risk of false positives (O’Reilly et al. 2008; Ferrer-Admetlla et al. 2014).
The observed relative EHH can finally be compared with a distribution of simulated
neutral relative EHH. Sabeti et al. (2002) successfully used the EHH test to detect
strong partial sweeps at the G6PD and CD40L loci. To this date, the G6PD locus is
one of the best examples of an ongoing, partial sweep in the human genome.
4.2.2.2 Integrated Haplotype Score (iHS)
A major limitation of the EHH test is that it uses an arbitrary distance to the core
haplotype. One can imagine the curve of EHH as a function of the distance to the
core haplotype. The area under the curve captures more information about haplotype
structure and represents a less arbitrary statistic than EHH measured for a fixed
distance to the core haplotype. In this respect, the integrated Haplotype Score (iHS)
(Voight et al. 2006) is an improved, integrated, and standardized version of EHH.
Unlike the classic EHH, iHS is oriented based on the derived or ancestral status
of alleles. Single ancestral or derived alleles are used as core haplotypes, and iHS
