9 Natural Selection, Genetic Variation, and Human Diversity
213
Strikingly, studies of different pastoral populations have identified several distinct
causative lactase-persistence alleles in an upstream intron of the gene MCM6
(Enattah et al. 2002, 2008; Tishkoff et al. 2007; Ingram et al. 2009; Gallego Romero
et al. 2012; Jones et al. 2013), which shows enhancer activity (Olds and Sibley
2003); thus, there has been convergent adaptive evolution influencing transcriptional
regulation of the LCT gene.
Another well-documented example of positive selection is found at the gene
DARC, which encodes the Duffy antigen receptor for chemokines. The malaria
parasite Plasmodium vivax requires the DARC protein to infect red blood cells,
but the FY*O allele prevents expression of DARC exclusively in the red blood
cells and therefore provides malaria resistance (Harris and Meyer 2006). FY*O is
essentially fixed in sub-Saharan Africa, where P. vivax was likely a strong selective
pressure in the past, but it is found at very low frequency in non-Africans. Phenotype
and population frequency information support the hypothesis that positive selection
drove FY*O to fixation in Africans, while non-Africans experienced no such
selection (Hamblin et al. 2002). In concordance with this hypothesis, FY*O has
a very high level of differentiation from non-Africans as measured by F ST , when
compared to the FY*A and FY*B alleles (Hamblin et al. 2002). The DARC region
also shows a skew toward rare variants (positive Tajima’s D) and an excess of
high-frequency derived variants (Hamblin et al. 2002). Notably, DARC shows
significantly lower variation within Africans than in non-Africans, which is opposite
the pattern expected given our knowledge of human migration history (Hamblin et
al. 2002).
Although other success stories arising from candidate gene studies exist, such
approaches of analyzing loci individually for signatures of selection are fraught with
difficulty. Most importantly, interpreting patterns of genetic variation for a single
locus at a time is difficult because of the confounding influence that demographic
history imparts on patterns of DNA sequence variation (Akey et al. 2004; Stajich
and Hahn 2005; Akey 2009; Li et al. 2012). As technologies became available to
more comprehensively survey patterns of human genetic variation in large sample
sizes, candidate gene studies of selection gave way to genome-wide studies, which
we discuss below.
9.4.2 Genome-Wide Scans for Selection
9.4.2.1 Insights from Genome-Wide Scans
Genome-wide scans provide a more comprehensive and unbiased approach to
systematically search the genome for substrates of adaptive evolution and also
enable more sophisticated approaches to disentangle the confounding effects of
selection and demographic history. The first genome-wide scans for selection in
humans were made possible by the development of dense genome-wide SNP
genotype data and showed that selection footprint could be detected in such data
(Akey et al. 2002). Genome-wide scans have proliferated, using a wide variety
of statistical approaches and methods designed to identify genomic regions with
213
Strikingly, studies of different pastoral populations have identified several distinct
causative lactase-persistence alleles in an upstream intron of the gene MCM6
(Enattah et al. 2002, 2008; Tishkoff et al. 2007; Ingram et al. 2009; Gallego Romero
et al. 2012; Jones et al. 2013), which shows enhancer activity (Olds and Sibley
2003); thus, there has been convergent adaptive evolution influencing transcriptional
regulation of the LCT gene.
Another well-documented example of positive selection is found at the gene
DARC, which encodes the Duffy antigen receptor for chemokines. The malaria
parasite Plasmodium vivax requires the DARC protein to infect red blood cells,
but the FY*O allele prevents expression of DARC exclusively in the red blood
cells and therefore provides malaria resistance (Harris and Meyer 2006). FY*O is
essentially fixed in sub-Saharan Africa, where P. vivax was likely a strong selective
pressure in the past, but it is found at very low frequency in non-Africans. Phenotype
and population frequency information support the hypothesis that positive selection
drove FY*O to fixation in Africans, while non-Africans experienced no such
selection (Hamblin et al. 2002). In concordance with this hypothesis, FY*O has
a very high level of differentiation from non-Africans as measured by F ST , when
compared to the FY*A and FY*B alleles (Hamblin et al. 2002). The DARC region
also shows a skew toward rare variants (positive Tajima’s D) and an excess of
high-frequency derived variants (Hamblin et al. 2002). Notably, DARC shows
significantly lower variation within Africans than in non-Africans, which is opposite
the pattern expected given our knowledge of human migration history (Hamblin et
al. 2002).
Although other success stories arising from candidate gene studies exist, such
approaches of analyzing loci individually for signatures of selection are fraught with
difficulty. Most importantly, interpreting patterns of genetic variation for a single
locus at a time is difficult because of the confounding influence that demographic
history imparts on patterns of DNA sequence variation (Akey et al. 2004; Stajich
and Hahn 2005; Akey 2009; Li et al. 2012). As technologies became available to
more comprehensively survey patterns of human genetic variation in large sample
sizes, candidate gene studies of selection gave way to genome-wide studies, which
we discuss below.
9.4.2 Genome-Wide Scans for Selection
9.4.2.1 Insights from Genome-Wide Scans
Genome-wide scans provide a more comprehensive and unbiased approach to
systematically search the genome for substrates of adaptive evolution and also
enable more sophisticated approaches to disentangle the confounding effects of
selection and demographic history. The first genome-wide scans for selection in
humans were made possible by the development of dense genome-wide SNP
genotype data and showed that selection footprint could be detected in such data
(Akey et al. 2002). Genome-wide scans have proliferated, using a wide variety
of statistical approaches and methods designed to identify genomic regions with
