104
R. E. Graff et al.
5.5.3 Analysis of Common Variants
The conventional analysis plan for genome-wide data is a series of statistical tests
that examine each locus independently for an association with the phenotype of
interest. In the case of binary phenotypes, the simplest approach to these tests is
contingency tables of counts of disease status by either genotype or allele count
(Clarke et al. 2011). One can use a series of χ 2 or Cochran-Armitage trend tests to
evaluate the independence of the rows and columns of each table. More commonly
used than contingency tables is a regression approach, linear for quantitative traits
and logistic for case-control traits, with categorical predictor variables for the
genotypes. Regression models are generally favored because they allow adjustment
for covariates, such as principal components (PC) of genetic ancestry.
For quantitative phenotypes, one can use a linear regression model framework,
y = α + xβ + cγ, to model association with phenotype, where x is a matrix (or
vector) of genotypes, c is a matrix of covariates such as ancestral PCs, and β and
γ are the corresponding vectors of regression coefficients. For binary phenotypes,
one would implement a logit link function. Regardless of the link function for the
outcome, the βs are the parameters of interest, and we can test the null hypothesis
of no association between x and y, H 0 : β = 0.
In addition to maximum likelihood-based regression models, linear mixed
models (LMM) have become increasingly popular in GWAS, motivated by the
computational challenges of analyzing datasets with large numbers of subjects and
genetic variants. One of the most attractive features of LMM is their ability to
control for confounding due to population structure by directly modeling relatedness
among individuals, thereby improving power relative to standard GWAS with
adjustment for PCs (Yang et al. 2014; Zhang et al. 2010; Zhou and Stephens
2012). The most recent addition to this class of methods, BOLT-LMM, adopts a
Bayesian perspective by imposing a prior distribution on SNP effect sizes. It does
not require computing or storing a genetic relationship matrix, which substantially
reduces computational time compared to other methods (Loh et al. 2015). However,
applying BOLT-LMM to case-control data can be problematic since genetic effects
are estimated on the observed 0–1 scale rather than the odds ratio scale. As a
result, transformations are required to make LMM-based results for binary traits
comparable with logistic regression (Lloyd-Jones et al. 2018).
Regardless of the structure of the phenotypic data, there are several ways in
which one might code the genotype data for association tests; the choice made
should reflect the assumed mode of inheritance and genetic effect. In GWAS, the
genotypes are usually coded as 0, 1, or 2 to reflect the number of effect alleles.
This coding assumes that each additional copy of the variant allele increases the
phenotype or log risk of disease by the same amount. The approach is fairly robust
to incorrect assumptions about the mode of inheritance and has reasonable power to
detect both additive and dominant genetic effects. It may, however, be underpowered
if the true mode of inheritance is recessive (Lettre et al. 2007). If one believes the
mode of inheritance to be recessive, then one may use an alternative genetic model
R. E. Graff et al.
5.5.3 Analysis of Common Variants
The conventional analysis plan for genome-wide data is a series of statistical tests
that examine each locus independently for an association with the phenotype of
interest. In the case of binary phenotypes, the simplest approach to these tests is
contingency tables of counts of disease status by either genotype or allele count
(Clarke et al. 2011). One can use a series of χ 2 or Cochran-Armitage trend tests to
evaluate the independence of the rows and columns of each table. More commonly
used than contingency tables is a regression approach, linear for quantitative traits
and logistic for case-control traits, with categorical predictor variables for the
genotypes. Regression models are generally favored because they allow adjustment
for covariates, such as principal components (PC) of genetic ancestry.
For quantitative phenotypes, one can use a linear regression model framework,
y = α + xβ + cγ, to model association with phenotype, where x is a matrix (or
vector) of genotypes, c is a matrix of covariates such as ancestral PCs, and β and
γ are the corresponding vectors of regression coefficients. For binary phenotypes,
one would implement a logit link function. Regardless of the link function for the
outcome, the βs are the parameters of interest, and we can test the null hypothesis
of no association between x and y, H 0 : β = 0.
In addition to maximum likelihood-based regression models, linear mixed
models (LMM) have become increasingly popular in GWAS, motivated by the
computational challenges of analyzing datasets with large numbers of subjects and
genetic variants. One of the most attractive features of LMM is their ability to
control for confounding due to population structure by directly modeling relatedness
among individuals, thereby improving power relative to standard GWAS with
adjustment for PCs (Yang et al. 2014; Zhang et al. 2010; Zhou and Stephens
2012). The most recent addition to this class of methods, BOLT-LMM, adopts a
Bayesian perspective by imposing a prior distribution on SNP effect sizes. It does
not require computing or storing a genetic relationship matrix, which substantially
reduces computational time compared to other methods (Loh et al. 2015). However,
applying BOLT-LMM to case-control data can be problematic since genetic effects
are estimated on the observed 0–1 scale rather than the odds ratio scale. As a
result, transformations are required to make LMM-based results for binary traits
comparable with logistic regression (Lloyd-Jones et al. 2018).
Regardless of the structure of the phenotypic data, there are several ways in
which one might code the genotype data for association tests; the choice made
should reflect the assumed mode of inheritance and genetic effect. In GWAS, the
genotypes are usually coded as 0, 1, or 2 to reflect the number of effect alleles.
This coding assumes that each additional copy of the variant allele increases the
phenotype or log risk of disease by the same amount. The approach is fairly robust
to incorrect assumptions about the mode of inheritance and has reasonable power to
detect both additive and dominant genetic effects. It may, however, be underpowered
if the true mode of inheritance is recessive (Lettre et al. 2007). If one believes the
mode of inheritance to be recessive, then one may use an alternative genetic model
