90
R. E. Graff et al.
utilized family structures ranging from sibling pairs to large multiplex pedigrees.
Such studies use families with numerous disease-affected individuals to evaluate
markers spaced widely across the genome, at intervals of up to 20 million base pairs,
and to examine how these markers segregate with the disease phenotype across
multiple families (Botstein et al. 1980). Linkage analyses are often successful in
the evaluation of rare and/or monogenetic disorders but are generally underpowered
to detect genetic factors with subtle effects on complex diseases. They also have low
resolution on account of the limited number of meioses from one generation to the
next within families (Risch and Merikangas 1996).
Given that high-penetrance genes co-segregating in affected families have turned
out to be relatively rare, association studies have become the far more common and
more powerful tool to investigate genetic relationships (Claussnitzer et al. 2020).
They rely on historical recombination events from millions of years of evolution
and thus do not require pedigree information or controlled crosses to identify genetic
variants associated with the phenotype. In addition, because most association studies
leverage the phenomenon of linkage disequilibrium (LD) to localize such variants,
they can detect causal loci within narrower regions and allow for genetic mapping
at a finer scale than linkage studies (Xiong and Guo 1997). That is, association
studies do not require the direct evaluation of postulated causal variants. Rather, they
may utilize LD to indirectly evaluate genetic variants neighboring those assayed
(see Chap. 2 on LD). Moreover, genome-wide association studies (GWAS) allow
investigators to broadly search the genome for disease-causing variants in a manner
that is relatively agnostic to previous biological knowledge.
The fundamental approach to any genetic association study is based on the
following premise: compare the frequency of the genetic characteristic of interest
across individuals with different values for the phenotype of interest. Consider,
for example, a single-nucleotide polymorphism (SNP) with effect allele A under
investigation in a standard analysis of a binary phenotype (Fig. 5.1). To determine
whether or not the SNP is associated with the phenotype, one would calculate the
frequency of the effect allele in cases and controls. When the frequency is greater
in individuals with the phenotype than in those without it, then the effect allele is
positively associated with the phenotype (as in the figure). When the opposite is
true, then the effect allele is inversely associated. In GWAS, these associations are
estimated for every SNP measured across the entire genome.
Genomic research traverses genetic sequence information, protein products, and
the eventual expression of traits. It may also utilize a range of organisms; only
one facet is the study of humans. Our focus in this chapter is on population-based
genetic association studies in humans, in which data are derived from unrelated
individuals. Relative to family-based association studies, population-based studies
are the more common—and often more powerful—approach to the evaluation
of genetic associations. In describing types of association studies (Sect. 5.2),
considerations in their design (Sect. 5.3), measurement of genetic information (Sect.
5.4), and analytical techniques (Sect. 5.5), we aim to provide a basis on which
readers can build their own efforts to characterize associations between genetic
polymorphisms and measured phenotypes.
R. E. Graff et al.
utilized family structures ranging from sibling pairs to large multiplex pedigrees.
Such studies use families with numerous disease-affected individuals to evaluate
markers spaced widely across the genome, at intervals of up to 20 million base pairs,
and to examine how these markers segregate with the disease phenotype across
multiple families (Botstein et al. 1980). Linkage analyses are often successful in
the evaluation of rare and/or monogenetic disorders but are generally underpowered
to detect genetic factors with subtle effects on complex diseases. They also have low
resolution on account of the limited number of meioses from one generation to the
next within families (Risch and Merikangas 1996).
Given that high-penetrance genes co-segregating in affected families have turned
out to be relatively rare, association studies have become the far more common and
more powerful tool to investigate genetic relationships (Claussnitzer et al. 2020).
They rely on historical recombination events from millions of years of evolution
and thus do not require pedigree information or controlled crosses to identify genetic
variants associated with the phenotype. In addition, because most association studies
leverage the phenomenon of linkage disequilibrium (LD) to localize such variants,
they can detect causal loci within narrower regions and allow for genetic mapping
at a finer scale than linkage studies (Xiong and Guo 1997). That is, association
studies do not require the direct evaluation of postulated causal variants. Rather, they
may utilize LD to indirectly evaluate genetic variants neighboring those assayed
(see Chap. 2 on LD). Moreover, genome-wide association studies (GWAS) allow
investigators to broadly search the genome for disease-causing variants in a manner
that is relatively agnostic to previous biological knowledge.
The fundamental approach to any genetic association study is based on the
following premise: compare the frequency of the genetic characteristic of interest
across individuals with different values for the phenotype of interest. Consider,
for example, a single-nucleotide polymorphism (SNP) with effect allele A under
investigation in a standard analysis of a binary phenotype (Fig. 5.1). To determine
whether or not the SNP is associated with the phenotype, one would calculate the
frequency of the effect allele in cases and controls. When the frequency is greater
in individuals with the phenotype than in those without it, then the effect allele is
positively associated with the phenotype (as in the figure). When the opposite is
true, then the effect allele is inversely associated. In GWAS, these associations are
estimated for every SNP measured across the entire genome.
Genomic research traverses genetic sequence information, protein products, and
the eventual expression of traits. It may also utilize a range of organisms; only
one facet is the study of humans. Our focus in this chapter is on population-based
genetic association studies in humans, in which data are derived from unrelated
individuals. Relative to family-based association studies, population-based studies
are the more common—and often more powerful—approach to the evaluation
of genetic associations. In describing types of association studies (Sect. 5.2),
considerations in their design (Sect. 5.3), measurement of genetic information (Sect.
5.4), and analytical techniques (Sect. 5.5), we aim to provide a basis on which
readers can build their own efforts to characterize associations between genetic
polymorphisms and measured phenotypes.
