5 Methods for Association Studies
113
Under a null hypothesis of no true associations in a GWAS dataset, P values would
conform to a uniform distribution between zero and one. The false discovery rate
essentially corrects for the expected number of false discoveries under this null
distribution. While it typically allows researchers to reject more null hypotheses
than would a Bonferroni adjustment, it may still be overly rigorous in the context of
GWAS or large-scale candidate gene association studies. In such scenarios, one may
implement weighted or stratified false discovery rates to achieve greater power to
detect true associations in subsets of SNPs with a higher proportion of true positives
than in the full set of SNPs (Genovese et al. 2006; Greenwood et al. 2007; Roeder
et al. 2006).
5.5.8.4 Bayesian Approach
Rather than correct for multiple comparisons via traditional frequentist methods,
one might choose to implement a Bayesian approach to the false discovery rate.
Under a Bayesian framework, the Bayes factor quantifies the strength of evidence
for an association between a SNP and phenotype. It is weighed against the prior
probability of an association to arrive at the posterior probability (Stephens and
Balding 2009). The calculation does not reference the number of SNPs tested.
While the expected number of false-positive associations will increase as more
tests are performed, so too will the number of true positive associations under a
reasonable set of assumptions. As such, the ratio of true to false positives will remain
roughly constant. Several software packages (e.g., SNPTEST (Marchini et al. 2007)
and BIMBAM (Servin and Stephens 2007)) accommodate genome-wide Bayesian
analyses.
5.6
Concluding Remarks
Well-executed genetic association studies can contribute immensely to our understanding of the underpinnings of disease. For meaningful conclusions to be drawn,
it is critical that they be designed with an appropriate population consisting
of a sufficient number of subjects. This number will depend upon the design
selected—candidate gene, GWAS, or otherwise—and should take into account the
methodological nuances thereof. Accurate measurement of genetic information is
also integral to the success of an association study, as are the quality control
checks that validate it. Then, statistical analyses must consider the specific research
question at hand, so as to make decisions that will best answer it.
There remain many gaps in our understanding and many association studies
that have the potential to fill them in moving forward. Since the proposal of an
exposome in 2005 (Wild 2005), investigators have striven to conceive of methods
that incorporate all of the exposures that individuals experience in a lifetime into
the study of their genetics. They have also been busy considering the question
of pleiotropy so as to identify genes that affect multiple, sometimes seemingly
unrelated, phenotypes. Many are developing methods that will better address rare
variants and interactions. The collection of these efforts will further improve our
113
Under a null hypothesis of no true associations in a GWAS dataset, P values would
conform to a uniform distribution between zero and one. The false discovery rate
essentially corrects for the expected number of false discoveries under this null
distribution. While it typically allows researchers to reject more null hypotheses
than would a Bonferroni adjustment, it may still be overly rigorous in the context of
GWAS or large-scale candidate gene association studies. In such scenarios, one may
implement weighted or stratified false discovery rates to achieve greater power to
detect true associations in subsets of SNPs with a higher proportion of true positives
than in the full set of SNPs (Genovese et al. 2006; Greenwood et al. 2007; Roeder
et al. 2006).
5.5.8.4 Bayesian Approach
Rather than correct for multiple comparisons via traditional frequentist methods,
one might choose to implement a Bayesian approach to the false discovery rate.
Under a Bayesian framework, the Bayes factor quantifies the strength of evidence
for an association between a SNP and phenotype. It is weighed against the prior
probability of an association to arrive at the posterior probability (Stephens and
Balding 2009). The calculation does not reference the number of SNPs tested.
While the expected number of false-positive associations will increase as more
tests are performed, so too will the number of true positive associations under a
reasonable set of assumptions. As such, the ratio of true to false positives will remain
roughly constant. Several software packages (e.g., SNPTEST (Marchini et al. 2007)
and BIMBAM (Servin and Stephens 2007)) accommodate genome-wide Bayesian
analyses.
5.6
Concluding Remarks
Well-executed genetic association studies can contribute immensely to our understanding of the underpinnings of disease. For meaningful conclusions to be drawn,
it is critical that they be designed with an appropriate population consisting
of a sufficient number of subjects. This number will depend upon the design
selected—candidate gene, GWAS, or otherwise—and should take into account the
methodological nuances thereof. Accurate measurement of genetic information is
also integral to the success of an association study, as are the quality control
checks that validate it. Then, statistical analyses must consider the specific research
question at hand, so as to make decisions that will best answer it.
There remain many gaps in our understanding and many association studies
that have the potential to fill them in moving forward. Since the proposal of an
exposome in 2005 (Wild 2005), investigators have striven to conceive of methods
that incorporate all of the exposures that individuals experience in a lifetime into
the study of their genetics. They have also been busy considering the question
of pleiotropy so as to identify genes that affect multiple, sometimes seemingly
unrelated, phenotypes. Many are developing methods that will better address rare
variants and interactions. The collection of these efforts will further improve our
