162
B. F. Voight
than common ones (Pritchard 2001; Schork et al. 2009; Gibson 2012), thus motivating technologies designed to characterize this spectrum of alleles carefully from
selected patient populations (Cirulli and Goldstein 2010). Empirical observation
supports the model that interindividual risk to complex traits is influenced by a
spectrum of alleles that are rare and common and weak and highly penetrant and
that the relative magnitude and effect in this spectrum depends on many factors,
including the intrinsic complexity of the trait (Iyengar and Elston 2007) as well the
historical fitness consequences of the phenotype studied (Eyre-Walker 2010; Simons
et al. 2014; Lohmueller 2014).
7.2.1 Study Design, Quality Control, and the Search for Technical
Biases
One of the major contributions of the genome-wide association approach was
the development of rigorous approaches to the control of the quality of samples,
genotypes, and the fundamental praxis of careful evaluation of data for technical
or other biases that can induce excess false-positive associations. It is obvious
to state that the best study is one where every sample and all polymorphisms
are accurately genotyped at high fidelity. However, even in the best-designed
study, high-throughput genomics often trades off some accuracy (ideally with
errors distributed randomly) for enhanced information collection. Before the first
studies were performed, it was difficult to predict all sources of errors but more
importantly which type of error would routinely induce false positives if not
controlled. For example, some of the first genotype-calling methods (e.g., the
dynamic modeling algorithm (Di et al. 2005)) were, for some polymorphic sites,
biased against heterozygous genotype calls (Rabbee and Speed 2005). Miscalling
heterozygotes—either incorrectly or by not calling—at best can result in reduced
power for association or, at worst, could result in false-positive association if cases
and controls were unbalanced in genotype calling (Zeggini and Morris 2011).
Today, the field enjoys a number of highly accurate genotype-calling algorithms
paired with workflows that instantiate a number of quality checks for samples and
genotyping data, for example, checks for the effects of “batches” of samples, gender
misclassification, various tests for missing genotypes, examination of quantilequantile plots, etc. Interested readers can consult a number of areas where checks
have been enumerated in detail elsewhere (Neale and Purcell 2008; Weale 2010) and
codified into association testing software to facilitate quality evaluations for these
purposes (Purcell et al. 2007).
After many studies have been performed, one broad lesson that has been
understood now is that errors largely can be traced back to the study design,
the source of sample DNA, how samples are processed through the lab, the
technology used to generate data, and the algorithm used to perform genotyping.
This basic insight has then fed back into the study design and the protocol by
which genotyping is performed. In many cases, this has resulted in a closer working
relationship between the statisticians and epidemiologists who helped plan the
B. F. Voight
than common ones (Pritchard 2001; Schork et al. 2009; Gibson 2012), thus motivating technologies designed to characterize this spectrum of alleles carefully from
selected patient populations (Cirulli and Goldstein 2010). Empirical observation
supports the model that interindividual risk to complex traits is influenced by a
spectrum of alleles that are rare and common and weak and highly penetrant and
that the relative magnitude and effect in this spectrum depends on many factors,
including the intrinsic complexity of the trait (Iyengar and Elston 2007) as well the
historical fitness consequences of the phenotype studied (Eyre-Walker 2010; Simons
et al. 2014; Lohmueller 2014).
7.2.1 Study Design, Quality Control, and the Search for Technical
Biases
One of the major contributions of the genome-wide association approach was
the development of rigorous approaches to the control of the quality of samples,
genotypes, and the fundamental praxis of careful evaluation of data for technical
or other biases that can induce excess false-positive associations. It is obvious
to state that the best study is one where every sample and all polymorphisms
are accurately genotyped at high fidelity. However, even in the best-designed
study, high-throughput genomics often trades off some accuracy (ideally with
errors distributed randomly) for enhanced information collection. Before the first
studies were performed, it was difficult to predict all sources of errors but more
importantly which type of error would routinely induce false positives if not
controlled. For example, some of the first genotype-calling methods (e.g., the
dynamic modeling algorithm (Di et al. 2005)) were, for some polymorphic sites,
biased against heterozygous genotype calls (Rabbee and Speed 2005). Miscalling
heterozygotes—either incorrectly or by not calling—at best can result in reduced
power for association or, at worst, could result in false-positive association if cases
and controls were unbalanced in genotype calling (Zeggini and Morris 2011).
Today, the field enjoys a number of highly accurate genotype-calling algorithms
paired with workflows that instantiate a number of quality checks for samples and
genotyping data, for example, checks for the effects of “batches” of samples, gender
misclassification, various tests for missing genotypes, examination of quantilequantile plots, etc. Interested readers can consult a number of areas where checks
have been enumerated in detail elsewhere (Neale and Purcell 2008; Weale 2010) and
codified into association testing software to facilitate quality evaluations for these
purposes (Purcell et al. 2007).
After many studies have been performed, one broad lesson that has been
understood now is that errors largely can be traced back to the study design,
the source of sample DNA, how samples are processed through the lab, the
technology used to generate data, and the algorithm used to perform genotyping.
This basic insight has then fed back into the study design and the protocol by
which genotyping is performed. In many cases, this has resulted in a closer working
relationship between the statisticians and epidemiologists who helped plan the
