5 Methods for Association Studies
105
that assumes that two copies of the risk allele are necessary to result in phenotype
susceptibility. For such models, the heterozygote and wild-type homozygote are
collapsed into a single category, and the genotypic exposure is treated as binary.
The genotypic exposure is also treated as binary for models that assume a dominant
mode of inheritance, but the heterozygote and mutant homozygote are collapsed
separately from the wild-type homozygote. These dominant models assume that
a single copy of the risk allele is expected to result in phenotype susceptibility.
Both recessive and dominant models force heterozygotes to have the same risk or
mean phenotype as one of the homozygotes. If investigators do not have an a priori
hypothesis as to the mode of inheritance, they may choose to assess several different
genetic models. Doing so, however, requires additional corrections for multiple
testing. Another option when there is no a priori hypothesis would be to avoid any
assumptions about how the risk for heterozygotes compares with both homozygotes.
In such codominant models, maintaining the three distinct genotype classes requires
two degrees of freedom, while the other models require only one, thereby making
the latter more attractive if the genetic effect approximately follows one of their
modes of inheritance.
To visualize the results from analyses of common variants, particularly from
GWAS, investigators often generate Manhattan plots (e.g., Fig. 5.4). The x-axis of
these scatter plots is a chromosomal position, and the y-axis shows the P value for
association with the phenotype. Each point on the plot represents a single SNP, and
the height of each point depicts the strength of association between the SNP and the
phenotype. Manhattan plots with genome-wide significant results often exhibit clear
peaks where SNPs in LD show comparable signals. Those with points seemingly
scattered at random should be viewed with some skepticism.
Quantile-quantile (Q-Q) plots are another important visualization tool to evaluate
potential bias or quality control problems in GWAS results (e.g., Fig. 5.5). These
plots present the expected distribution of association test statistics for all SNPs
on the x-axis against the observed values on the y-axis. Deviation from the x = y
line suggests a systematic difference between cases and controls across the whole
genome (such as that which might occur in the presence of population stratification).
One should rather hope to see the plotted points fall on the x = y line until a curve
at the very end representing any true associations.
5.5.4 Analysis of Rare Variants
While many common variants that contribute to complex diseases have been
identified, the majority of variants contributing to disease susceptibility have yet to
be described. Rare variants, which are unlikely to be captured by GWAS focusing
on common SNPs, undoubtedly contribute to phenotype as well (Frazer et al. 2009;
Gorlov et al. 2008). Unfortunately, detecting associations between individual rare
variants and phenotypes can be difficult, even with large sample sizes; the low
frequency of rare variants in the population results in low power (Gorlov et al.
2008; Altshuler et al. 2008; Li and Leal 2008). To increase power, researchers have
105
that assumes that two copies of the risk allele are necessary to result in phenotype
susceptibility. For such models, the heterozygote and wild-type homozygote are
collapsed into a single category, and the genotypic exposure is treated as binary.
The genotypic exposure is also treated as binary for models that assume a dominant
mode of inheritance, but the heterozygote and mutant homozygote are collapsed
separately from the wild-type homozygote. These dominant models assume that
a single copy of the risk allele is expected to result in phenotype susceptibility.
Both recessive and dominant models force heterozygotes to have the same risk or
mean phenotype as one of the homozygotes. If investigators do not have an a priori
hypothesis as to the mode of inheritance, they may choose to assess several different
genetic models. Doing so, however, requires additional corrections for multiple
testing. Another option when there is no a priori hypothesis would be to avoid any
assumptions about how the risk for heterozygotes compares with both homozygotes.
In such codominant models, maintaining the three distinct genotype classes requires
two degrees of freedom, while the other models require only one, thereby making
the latter more attractive if the genetic effect approximately follows one of their
modes of inheritance.
To visualize the results from analyses of common variants, particularly from
GWAS, investigators often generate Manhattan plots (e.g., Fig. 5.4). The x-axis of
these scatter plots is a chromosomal position, and the y-axis shows the P value for
association with the phenotype. Each point on the plot represents a single SNP, and
the height of each point depicts the strength of association between the SNP and the
phenotype. Manhattan plots with genome-wide significant results often exhibit clear
peaks where SNPs in LD show comparable signals. Those with points seemingly
scattered at random should be viewed with some skepticism.
Quantile-quantile (Q-Q) plots are another important visualization tool to evaluate
potential bias or quality control problems in GWAS results (e.g., Fig. 5.5). These
plots present the expected distribution of association test statistics for all SNPs
on the x-axis against the observed values on the y-axis. Deviation from the x = y
line suggests a systematic difference between cases and controls across the whole
genome (such as that which might occur in the presence of population stratification).
One should rather hope to see the plotted points fall on the x = y line until a curve
at the very end representing any true associations.
5.5.4 Analysis of Rare Variants
While many common variants that contribute to complex diseases have been
identified, the majority of variants contributing to disease susceptibility have yet to
be described. Rare variants, which are unlikely to be captured by GWAS focusing
on common SNPs, undoubtedly contribute to phenotype as well (Frazer et al. 2009;
Gorlov et al. 2008). Unfortunately, detecting associations between individual rare
variants and phenotypes can be difficult, even with large sample sizes; the low
frequency of rare variants in the population results in low power (Gorlov et al.
2008; Altshuler et al. 2008; Li and Leal 2008). To increase power, researchers have
