110
R. E. Graff et al.
With the advent of high-throughput technology, investigators are exploring geneenvironment interactions at the genome-wide level. They are also realizing some
of the fundamental challenges of doing so. The genome-wide approach does not
make use of prior knowledge of biological processes and pathways. In addition, the
stringent significance threshold required due to the number of statistical tests may
preclude the identification of significant interactions.
Analysis approaches can focus on environmental interactions with single genes,
multiple genes, and/or biological pathways (Thomas 2010). Alternatively, one can
utilize available biomarker data that may reflect intermediate phenotypes to establish
informative priors in a hierarchical model framework (Li et al. 2012). Regardless, it
bears recognition that statistical interaction does not conclusively indicate causality.
Variants showing interaction may not be causal, and interactions may be significant
for reasons other than true association (as is true for any other type of association).
Still, the identification of gene-environment interactions with respect to disease risk
is of fundamental public health relevance.
5.5.7 Incorporating Covariates
5.5.7.1 Population Stratification
Population-based association studies are susceptible to a form of confounding
known as population stratification. It occurs whenever the gene of interest shows
pronounced variation in allele frequency across ancestral subgroups of the population, and these subgroups differ in their baseline risk of disease. In extreme
scenarios, population stratification can result from cryptic relatedness, wherein
some individuals in an ostensibly unrelated population are actually related.
In a sample with population stratification in any form, SNPs with large allele
frequency differences across groups are likely to be associated with the trait under
study. The first step toward dealing with the bias is to ensure that cases and controls
are well-matched in the study design phase. One can then evaluate the extent of
residual population stratification via Q–Q plots and their associated inflation factor,
lambda (λ). The latter is defined as the ratio of the median of the observed test
statistics relative to the expected median and reflects the excess false-positive rate.
When the value of lambda is inflated, one can adjust the test statistics by dividing
them by lambda, thereby reducing them, and then recalculating the associated P
values (Devlin and Roeder 1999).
In recent years, the more common approach to the management of population
stratification has been the measurement of the ancestry of each sample in the dataset
using PC methods (Price et al. 2006; Falush et al. 2003). These methods cluster
individuals together based on their ancestral populations, often by comparing them
with an external reference population such as the 1000 Genomes Project or TOPMed
imputation reference panel. With the results, one can then exclude samples that are
extremely different from the main clusters of individuals and then include the top
10 or so PCs as covariates in association analyses. A criticism of PCs is that they
are unable to differentiate between true signal due to polygenicity and confounding
R. E. Graff et al.
With the advent of high-throughput technology, investigators are exploring geneenvironment interactions at the genome-wide level. They are also realizing some
of the fundamental challenges of doing so. The genome-wide approach does not
make use of prior knowledge of biological processes and pathways. In addition, the
stringent significance threshold required due to the number of statistical tests may
preclude the identification of significant interactions.
Analysis approaches can focus on environmental interactions with single genes,
multiple genes, and/or biological pathways (Thomas 2010). Alternatively, one can
utilize available biomarker data that may reflect intermediate phenotypes to establish
informative priors in a hierarchical model framework (Li et al. 2012). Regardless, it
bears recognition that statistical interaction does not conclusively indicate causality.
Variants showing interaction may not be causal, and interactions may be significant
for reasons other than true association (as is true for any other type of association).
Still, the identification of gene-environment interactions with respect to disease risk
is of fundamental public health relevance.
5.5.7 Incorporating Covariates
5.5.7.1 Population Stratification
Population-based association studies are susceptible to a form of confounding
known as population stratification. It occurs whenever the gene of interest shows
pronounced variation in allele frequency across ancestral subgroups of the population, and these subgroups differ in their baseline risk of disease. In extreme
scenarios, population stratification can result from cryptic relatedness, wherein
some individuals in an ostensibly unrelated population are actually related.
In a sample with population stratification in any form, SNPs with large allele
frequency differences across groups are likely to be associated with the trait under
study. The first step toward dealing with the bias is to ensure that cases and controls
are well-matched in the study design phase. One can then evaluate the extent of
residual population stratification via Q–Q plots and their associated inflation factor,
lambda (λ). The latter is defined as the ratio of the median of the observed test
statistics relative to the expected median and reflects the excess false-positive rate.
When the value of lambda is inflated, one can adjust the test statistics by dividing
them by lambda, thereby reducing them, and then recalculating the associated P
values (Devlin and Roeder 1999).
In recent years, the more common approach to the management of population
stratification has been the measurement of the ancestry of each sample in the dataset
using PC methods (Price et al. 2006; Falush et al. 2003). These methods cluster
individuals together based on their ancestral populations, often by comparing them
with an external reference population such as the 1000 Genomes Project or TOPMed
imputation reference panel. With the results, one can then exclude samples that are
extremely different from the main clusters of individuals and then include the top
10 or so PCs as covariates in association analyses. A criticism of PCs is that they
are unable to differentiate between true signal due to polygenicity and confounding
