214
L. S. Emery and J. M. Akey
unusual patterns of variation that may be indicative of selection (Table 9.1). These
test statistics examine features including levels of population structure (Akey et al.
2002; Chen et al. 2010); amount and patterns of linkage disequilibrium (Sabeti et al.
2002, 2007; Voight et al. 2006); measures of the site frequency spectrum such as an
excess of high frequency derived alleles or an excess of rare variants (Tajima 1989;
Fay and Wu 2000; Grossman et al. 2010; Zhong et al. 2010, 2011); and differences
between population- and pedigree-based recombination rate estimates (O’Reilly et
al. 2008). It is important to note that no single test statistic is best suited for detecting
all types of selection and their sensitivity and specificity vary widely (Ronald and
Akey 2005; Biswas and Akey 2006; Sabeti et al. 2007). Additionally, some of the
methods that have been developed are appropriate for identifying selective sweeps
that have gone to completion (fixation of the causal adaptive allele), whereas others
are most appropriate for ongoing or incomplete sweeps.
Although there are still many significant advances to be made, the dozens of
published genome-wide scans for selection have provided a comprehensive firstpass picture of how, when, and where positive selection has affected the human
genome. Even the earliest of studies were able to show that the signatures of natural
selection are fairly evenly distributed throughout the genome rather than being
found in discrete clusters (Akey et al. 2002; International HapMap Consortium
2005; Nielsen et al. 2005). Additionally, scans have thus far identified about 10%
of the genome as being under recent positive selection within humans (Akey
2009). Furthermore, the number of regions identified as the targets of selection is
substantial, whether testing for the signatures of complete or incomplete sweeps
(Akey 2009).
In studies that have truly examined the whole genome, intergenic regions exhibit
a surprising number of selection signatures (Akey 2009). Although some of these
signals can be attributed to regulatory variants associated with nearby genes, some
are so far from the nearest gene that other interesting explanations are possible. The
functional substrate of selection in these regions could be regulatory elements, such
as distant enhancers. They could also be attributed to functional non-coding RNAs,
which we still know relatively little about. Another intriguing possibility is that
some intergenic targets of selection are structural elements responsible for genome
organization. For example, Williamson and colleagues observed many signals
indicating positive selection acting on centromeres. These centromeric regions could
be selected due to meiotic drive, since any variant that makes a chromosome more
likely to end up in the oocyte than in a polar body during female meiosis will have
a very strong selective advantage (Williamson et al. 2007). While it is evident that
the targets of recent positive selection are not restricted to genic or gene-associated
regions, these intergenic selected regions have been largely overlooked thus far and
warrant further study.
A majority of genome-wide scans have either limited their analyses to regions
around genes or focused on identified targets within or near genes. From these
results, we can confidently say that a wide variety of genes in the human genome
have been subject to recent positive selection. Some of these genes are also identified
as the targets of older selective events, such as those detected by interspecific
L. S. Emery and J. M. Akey
unusual patterns of variation that may be indicative of selection (Table 9.1). These
test statistics examine features including levels of population structure (Akey et al.
2002; Chen et al. 2010); amount and patterns of linkage disequilibrium (Sabeti et al.
2002, 2007; Voight et al. 2006); measures of the site frequency spectrum such as an
excess of high frequency derived alleles or an excess of rare variants (Tajima 1989;
Fay and Wu 2000; Grossman et al. 2010; Zhong et al. 2010, 2011); and differences
between population- and pedigree-based recombination rate estimates (O’Reilly et
al. 2008). It is important to note that no single test statistic is best suited for detecting
all types of selection and their sensitivity and specificity vary widely (Ronald and
Akey 2005; Biswas and Akey 2006; Sabeti et al. 2007). Additionally, some of the
methods that have been developed are appropriate for identifying selective sweeps
that have gone to completion (fixation of the causal adaptive allele), whereas others
are most appropriate for ongoing or incomplete sweeps.
Although there are still many significant advances to be made, the dozens of
published genome-wide scans for selection have provided a comprehensive firstpass picture of how, when, and where positive selection has affected the human
genome. Even the earliest of studies were able to show that the signatures of natural
selection are fairly evenly distributed throughout the genome rather than being
found in discrete clusters (Akey et al. 2002; International HapMap Consortium
2005; Nielsen et al. 2005). Additionally, scans have thus far identified about 10%
of the genome as being under recent positive selection within humans (Akey
2009). Furthermore, the number of regions identified as the targets of selection is
substantial, whether testing for the signatures of complete or incomplete sweeps
(Akey 2009).
In studies that have truly examined the whole genome, intergenic regions exhibit
a surprising number of selection signatures (Akey 2009). Although some of these
signals can be attributed to regulatory variants associated with nearby genes, some
are so far from the nearest gene that other interesting explanations are possible. The
functional substrate of selection in these regions could be regulatory elements, such
as distant enhancers. They could also be attributed to functional non-coding RNAs,
which we still know relatively little about. Another intriguing possibility is that
some intergenic targets of selection are structural elements responsible for genome
organization. For example, Williamson and colleagues observed many signals
indicating positive selection acting on centromeres. These centromeric regions could
be selected due to meiotic drive, since any variant that makes a chromosome more
likely to end up in the oocyte than in a polar body during female meiosis will have
a very strong selective advantage (Williamson et al. 2007). While it is evident that
the targets of recent positive selection are not restricted to genic or gene-associated
regions, these intergenic selected regions have been largely overlooked thus far and
warrant further study.
A majority of genome-wide scans have either limited their analyses to regions
around genes or focused on identified targets within or near genes. From these
results, we can confidently say that a wide variety of genes in the human genome
have been subject to recent positive selection. Some of these genes are also identified
as the targets of older selective events, such as those detected by interspecific
