108
R. E. Graff et al.
sets with coordinated expression changes that would not be detected by single
variant methods (e.g., testing one SNP at a time).
Gene set analysis generally consists of four steps: (1) determining the gene
sets to be tested, (2) selecting an appropriate set of hypotheses, (3) carrying out
corresponding statistical tests, and (4) evaluating the statistical significance of said
tests. Regarding step 2, there are three standard null hypotheses against which
investigators most often test (Dinu et al. 2009; Nam and Kim 2008; Tian et al.
2005). The first is the competitive null hypothesis, which states that the genes in
a set have the same association with phenotype as genes in the rest of the genome.
The second is the self-contained null hypothesis, which asserts that the genes in a
set are not associated with the disease phenotype. The third option, the mixed null
hypothesis, declares that none of the sets under consideration is associated with the
disease.
The set of hypotheses selected largely informs the tests that should be used for
analysis. To obtain a test statistic for the competitive null, a measure of association
should first be computed for each gene and the phenotype of interest. For genes
in a given set, the association measures should then be combined. To evaluate the
statistical significance of the combined test statistic, it should be compared against
the distribution under the null hypothesis, obtained by permuting the association
measures (Tian et al. 2005). The procedure is similar to obtaining a test statistic for
the self-contained null hypothesis, but the null distribution should be generated by
permuting the phenotypes across samples (Tian et al. 2005). Regardless of the test
statistic, larger magnitudes indicate increasing significance, and the sign indicates
the direction of change in phenotype.
The gene sets that are deemed significant are likely to depend on the choice
of methods implemented to analyze gene set associations (Elbers et al. 2009a, b).
Oftentimes, gene set analyses lack sufficient statistical power to detect gene sets
consisting of markers only weakly associated with disease, and they are prone to
several sources of bias, among which are gene set size, LD patterns, and overlapping
genes (Elbers et al. 2009b; Cantor et al. 2010; Hong et al. 2009; Wang et al. 2011;
Sun et al. 2019). It is important to consider and address all of these limitations when
interpreting results from gene set analyses.
5.5.5.2 Hierarchical Modeling
Hierarchical modeling leverages the abundance of bioinformatic data characterizing
the structural and functional roles of common variants analyzed for GWAS (Cantor
et al. 2010; Wang et al. 2010). It aims to incorporate a priori biological knowledge
via Bayesian methods (Cardin et al. 2012), stabilize effect and variance estimation
of SNP associations (Aragaki et al. 1997; Evangelou et al. 2014), and improve the
selection of SNPs for further evaluation (Witte 1997; Witte and Greenland 1996).
It also addresses issues of multiple comparisons in analyses of GWAS. Rather than
perform traditional single-locus analyses, hierarchical models output “knowledgebased” estimates of SNP effects, thereby improving the ranking of results from
GWAS.
Précédent

- 113/236

Suivant