5 Methods for Association Studies
109
The hierarchical modeling approach uses higher-level “priors” to model the
parameters of interest as random variables with a joint distribution that is a
function of hyperparameters (Witte 1997). In addition to information about x and
y (as defined above), one also utilizes information about similarities among the
components of β. For example, one might assume that associations corresponding
to markers that are located near one another on a particular chromosome might
be similar. Conditional on this additional information, one may fit a second-stage
generalized model for the expectation of β: f 2 (β | Z) = δ + Zπ. According to this
model, f 2 is a strictly increasing link function, and Z is a second-stage design matrix
expressing the similarities among the β. Hierarchical (i.e., posterior) estimates are
obtained by combining results from the different level models (Witte 1997).
5.5.6 Interactions
GWAS present an opportunity to go beyond single-locus analyses and into the
realm of gene-gene interactions throughout the genome. Given the number of
SNPs generally evaluated in GWAS, it would prove intractable to evaluate all
pairwise combinations. Instead, one can reduce the set of SNPs to further investigate
via one of several methods (McAllister et al. 2017). The first is to select an
arbitrary significance threshold for the set of single-locus analyses. One can then
evaluate all pairwise interactions between SNPs falling below the threshold, or
between such SNPs and all other SNPs. Implementing this method, however, will
preclude the discovery of combinations of markers that affect a significant change
in disease risk even when the individual markers’ marginal effects are statistically
undetectable. An alternative approach is to restrict the analysis of interactions to
SNPs with an established biological function or within a particular protein family.
A general comment regarding all analyses of interaction is that the scale (additive
or multiplicative) on which they are evaluated will impact the results.
The evaluation of gene-gene interactions is not limited to model-based methods. Multifactor dimensionality reduction (MDR) was developed to reduce the
dimensionality of multilocus data so as to improve the ability to detect gene-gene
interactions. MDR pools genotypes into high-risk and low-risk groups, thereby
reducing data to a single dimension. The method is nonparametric and model-free—
one need not make hypotheses regarding the values of any parameters or assume any
particular mode of inheritance (Motsinger and Ritchie 2006). The details of MDR
analyses have been well described (Hahn et al. 2003; Ritchie et al. 2001, 2003).
The study of gene-environment interactions is another critical component of
understanding the biological mechanisms of complex disease, heterogeneity across
studies, and susceptible subpopulations (Dick et al. 2015). Until recently, geneenvironment interaction studies have been largely carried out using candidate
approaches. Such studies require the identification of genes with related biological
functionality as well as knowledge of the mode of action through which environmental factors affect the genes of interest (Rava et al. 2013).
109
The hierarchical modeling approach uses higher-level “priors” to model the
parameters of interest as random variables with a joint distribution that is a
function of hyperparameters (Witte 1997). In addition to information about x and
y (as defined above), one also utilizes information about similarities among the
components of β. For example, one might assume that associations corresponding
to markers that are located near one another on a particular chromosome might
be similar. Conditional on this additional information, one may fit a second-stage
generalized model for the expectation of β: f 2 (β | Z) = δ + Zπ. According to this
model, f 2 is a strictly increasing link function, and Z is a second-stage design matrix
expressing the similarities among the β. Hierarchical (i.e., posterior) estimates are
obtained by combining results from the different level models (Witte 1997).
5.5.6 Interactions
GWAS present an opportunity to go beyond single-locus analyses and into the
realm of gene-gene interactions throughout the genome. Given the number of
SNPs generally evaluated in GWAS, it would prove intractable to evaluate all
pairwise combinations. Instead, one can reduce the set of SNPs to further investigate
via one of several methods (McAllister et al. 2017). The first is to select an
arbitrary significance threshold for the set of single-locus analyses. One can then
evaluate all pairwise interactions between SNPs falling below the threshold, or
between such SNPs and all other SNPs. Implementing this method, however, will
preclude the discovery of combinations of markers that affect a significant change
in disease risk even when the individual markers’ marginal effects are statistically
undetectable. An alternative approach is to restrict the analysis of interactions to
SNPs with an established biological function or within a particular protein family.
A general comment regarding all analyses of interaction is that the scale (additive
or multiplicative) on which they are evaluated will impact the results.
The evaluation of gene-gene interactions is not limited to model-based methods. Multifactor dimensionality reduction (MDR) was developed to reduce the
dimensionality of multilocus data so as to improve the ability to detect gene-gene
interactions. MDR pools genotypes into high-risk and low-risk groups, thereby
reducing data to a single dimension. The method is nonparametric and model-free—
one need not make hypotheses regarding the values of any parameters or assume any
particular mode of inheritance (Motsinger and Ritchie 2006). The details of MDR
analyses have been well described (Hahn et al. 2003; Ritchie et al. 2001, 2003).
The study of gene-environment interactions is another critical component of
understanding the biological mechanisms of complex disease, heterogeneity across
studies, and susceptible subpopulations (Dick et al. 2015). Until recently, geneenvironment interaction studies have been largely carried out using candidate
approaches. Such studies require the identification of genes with related biological
functionality as well as knowledge of the mode of action through which environmental factors affect the genes of interest (Rava et al. 2013).
