7 What Have We Learned from GWAS?
171
of alleles that underlie them. For many traits with association that are known,
this distribution across a range of phenotypes is, with occasional outliers, trending
toward “U-shaped,” where the effect size increases as risk alleles become infrequent
(e.g., Speliotes et al. (2010), Fig. 1c). However, this distribution is largely biased
due to the power of the study to discover effects, given sample sizes collected. A
more formal way to present this discussion is to formally test hypotheses about
frequencies and effect sizes, building confidence sets of models of both parameters
that fail to be rejected by the observed distribution of association scores in metaanalysis. As an example, power analysis focused on type 2 diabetes (and the sample
sizes collected to date), assuming a simple additive genetic model, and across
the range of “surveyable” genome, we can easily reject genetic models which
postulate that additional common variants (∼20–80% frequency) with modest effect
or greater (odds ratio of 1.2 or greater) remain to be discovered (Park et al. 2010).
While it might be tempting to model the distribution of effect sizes and allele
frequencies underlying complex disease, it is important to note that, except in a
few rare circumstances, the actual causal alleles are not known with precision.
More than likely, lead associations at common variation are simply strongly linked
with variation that is indeed causal. An alternative hypothesis is that common
variation is actually linked with rare, casual, coding mutations (potentially at
very long distances), and thus the observed association at common sites with
traits is merely “synthetic” (Dickson et al. 2010). While this hypothesis has some
theoretical hurdles to overcome in explaining the majority of common variant
associations (Wray et al. 2011; Anderson et al. 2011), empirical genetic studies are
slowly indicating that low-frequency and rare variation cannot explain all common
variant association signals. Coupled with examples of mechanistic explanations for
complex traits (Musunuru et al. 2010), noncoding variation increasingly appears
implicated in the biology of complex diseases. Systematic answers along all of these
lines will require a comprehensive panel of variation across the frequency spectrum
surveyed in a large number of samples.
7.3.8 Observational Epidemiology and Genetics Need Not Always
Agree
Prior to initiating genetic studies, epidemiological and heritability studies are
first initiated to quantify (and demonstrate) that specific clinical or physiological
phenotypes segregate within families and are amenable to genetic dissection. As
such, numerous epidemiological studies on a range of traits have been performing
and continue to be performed. One feature that emerges from such studies, beyond
the estimates of heritability for traits, is that many of these traits are correlated
with one another. This can be due to not only how the traits are measured but
also the intrinsic biology of the traits. This observation has been exploited with
important practical effect, for example, in using plasma-lipid levels as biomarkers
to predict the incidence of heart attack (Emerging Risk Factors Collaboration 2009).
Other correlations across traits are also well-known: between systolic and diastolic
171
of alleles that underlie them. For many traits with association that are known,
this distribution across a range of phenotypes is, with occasional outliers, trending
toward “U-shaped,” where the effect size increases as risk alleles become infrequent
(e.g., Speliotes et al. (2010), Fig. 1c). However, this distribution is largely biased
due to the power of the study to discover effects, given sample sizes collected. A
more formal way to present this discussion is to formally test hypotheses about
frequencies and effect sizes, building confidence sets of models of both parameters
that fail to be rejected by the observed distribution of association scores in metaanalysis. As an example, power analysis focused on type 2 diabetes (and the sample
sizes collected to date), assuming a simple additive genetic model, and across
the range of “surveyable” genome, we can easily reject genetic models which
postulate that additional common variants (∼20–80% frequency) with modest effect
or greater (odds ratio of 1.2 or greater) remain to be discovered (Park et al. 2010).
While it might be tempting to model the distribution of effect sizes and allele
frequencies underlying complex disease, it is important to note that, except in a
few rare circumstances, the actual causal alleles are not known with precision.
More than likely, lead associations at common variation are simply strongly linked
with variation that is indeed causal. An alternative hypothesis is that common
variation is actually linked with rare, casual, coding mutations (potentially at
very long distances), and thus the observed association at common sites with
traits is merely “synthetic” (Dickson et al. 2010). While this hypothesis has some
theoretical hurdles to overcome in explaining the majority of common variant
associations (Wray et al. 2011; Anderson et al. 2011), empirical genetic studies are
slowly indicating that low-frequency and rare variation cannot explain all common
variant association signals. Coupled with examples of mechanistic explanations for
complex traits (Musunuru et al. 2010), noncoding variation increasingly appears
implicated in the biology of complex diseases. Systematic answers along all of these
lines will require a comprehensive panel of variation across the frequency spectrum
surveyed in a large number of samples.
7.3.8 Observational Epidemiology and Genetics Need Not Always
Agree
Prior to initiating genetic studies, epidemiological and heritability studies are
first initiated to quantify (and demonstrate) that specific clinical or physiological
phenotypes segregate within families and are amenable to genetic dissection. As
such, numerous epidemiological studies on a range of traits have been performing
and continue to be performed. One feature that emerges from such studies, beyond
the estimates of heritability for traits, is that many of these traits are correlated
with one another. This can be due to not only how the traits are measured but
also the intrinsic biology of the traits. This observation has been exploited with
important practical effect, for example, in using plasma-lipid levels as biomarkers
to predict the incidence of heart attack (Emerging Risk Factors Collaboration 2009).
Other correlations across traits are also well-known: between systolic and diastolic
