222
L. S. Emery and J. M. Akey
the strongest signals of recent selection are detectable. While it is likely that signals
of very weak selection will always remain beyond the limits of detection, the use
of the outlier approach and other methods without formal significance thresholds
makes it particularly difficult to sort out false positives and false negatives (Kelley
et al. 2006). Furthermore, our ability to detect a signature of selection remains
largely dependent upon levels of background linkage disequilibrium in the regions
(O’Reilly et al. 2008; Stephan 2010). The result is that targets of recent selection
are easier to identify when they are found in regions of unusually low-background
LD (O’Reilly et al. 2008; Akey 2009). This issue has been addressed by some of the
LD-based statistics, such as iHS and EHH, but more comprehensive methods should
also be developed (O’Reilly et al. 2008). The problems of detection are exacerbated
by conflicting results from various different neutrality test statistics. Since different
statistics measure different kinds of signatures of selection, it is not entirely clear
what combination of signals is most indicative of selection. Composite likelihood
methods have been developed to address this problem, and their further development
is warranted (Grossman et al. 2010).
Perhaps the most glaring drawback of existing genome-wide scans is the repeated
use of the same few population panel SNP data sets, including SeattleSNPs (http://
pga.gs.washington.edu), the SNP Consortium (Altshuler et al. 2000), the Perlegen
SNPs (Hinds et al. 2005), and the HapMap SNPs (phase 1 (International HapMap
Consortium 2005) or 2 (International HapMap Consortium et al. 2007)). Of the 31
studies presented in Table 9.1, only 8 (26%) examined other population data sets.
While the use of the same data sets makes comparisons between studies easier, it
also significantly reduces the applicability of those results to other data sets. Studies
examining new populations are likely to provide much more novel information
than yet another study on a panel of one African, one Asian, and one European
population. Another major drawback to the existing data sets is the pervasive
problem of ascertainment bias in SNP identification. SNP genotyping platforms
were developed using nonuniform SNP discovery techniques that were biased
toward identifying polymorphisms within European ancestry populations (Clark
et al. 2005). Therefore, the information these SNP chips provide about variation
in non-European populations is very difficult to interpret accurately. Furthermore,
SNP data is too sparse to allow the detection of the causative variant underlying
a signature of selection in most cases (Grossman et al. 2010). The use of wholegenome sequences for detecting selection will solve both problems of ascertainment
bias and resolution, as shown in recent studies using whole-exome sequences to
detect causative variants (Tennessen et al. 2010, 2012).
Finally, the usefulness of current genome-wide scans for recent selection is
severely limited by the simplicity of selection models they employ. The vast
majority of current genome-wide scans have been predicated on the model of
the classic selective sweep, in which a new mutation is immediately beneficial
upon its introduction into the population and quickly reaches fixation, sweeping
along linked neutral variants (Fig. 9.1a). While some attention has focused on
detecting signatures of incomplete or in-progress sweeps (Voight et al. 2006), that
is the only variation on the classic sweep that has been accounted for in genome-
Précédent

- 224/236

Suivant