5 Methods for Association Studies
103
Reference Haplotypes
(e.g., HapMap)
Observed Genotypes
.
.
.
.
.
.
.
.
G
G
.
.
.
.
.
.
.
.
.
.
.
.
.
.
G
C
.
.
.
.
.
.
.
.
G
A
.
.
.
.
.
.
C
C
C
C
C
T
C
C
C
C
G
G
C
G
G
G
G
G
G
G
A
A
A
A
A
G
A
A
A
A
G
G
A
A
G
G
G
G
G
A
A
G
A
G
A
A
A
A
A
G
T
T
C
C
C
T
T
C
C
C
C
C
T
T
T
C
C
T
T
T
T
T
C
C
C
T
T
C
C
C
C
C
T
T
T
C
C
T
T
T
C
C
T
T
C
C
C
T
C
T
T
C
T
T
C
C
C
T
C
T
T
G
T
T
G
G
G
T
G
T
C
G
C
A
A
A
A
C
A
C
T
C
T
T
C
C
C
T
C
T
T
C
T
T
C
C
C
T
C
T
C
T
C
C
T
T
T
T
T
C
T
C
T
T
T
C
T
T
C
T
G
G
G
G
G
A
G
G
G
G
T
T
T
T
T
T
T
T
T
T
G
G
G
G
G
G
G
A
G
G
C
G
C
C
C
G
C
C
C
C
Imputed Genotypes
c
c
g
g
a
a
g
a
G
G
t
c
c
t
t
c
c
t
c
t
c
t
g
t
G
C
c
t
c
t
t
t
c
c
G
A
t
t
g
g
g
g
Fig. 5.3 Schematic of genotype imputation modified from Li and colleagues (2009). Observed
genotypes are compared to haplotypes in a reference panel to fill in unobserved genotypes
5.5.2 Data Imputation
Above, we discussed the exploitation of LD to capture common variation in the
human genome without directly genotyping every SNP. Using LD patterns and
haplotype frequencies (e.g., from the 1000 Genomes Project or TOPMed imputation
reference panel), it is possible to impute data for SNPs that are not directly
genotyped (Fig. 5.3) (Li et al. 2009). First, directly genotyped markers are compared
to a variant-dense reference panel that contains haplotypes drawn from the same
population as the study sample. A collection of shared haplotypes is then identified,
and genotypes missing from the study panel can be inferred from the matching
reference haplotypes. Because the study sample may match multiple reference
haplotypes, one might opt to give a score or probability for an imputed marker rather
than a definitive allele. In such scenarios, uncertainty can be incorporated into the
analysis of imputed data, typically with Bayesian methods (Marchini et al. 2007).
A less computationally intensive method involves pre-phasing, in which haplotypes
are first estimated for every individual, followed by genotype imputation using the
reference panel for each haplotype. This method also makes it faster to execute the
imputation step with different reference panels as they become updated, since the
genotypes need only be phased once and the estimated haplotypes are saved for
future use (Howie et al. 2012). Note that imputation is especially useful for metaanalyzing results across studies that rely on different genotyping platforms.
Précédent

- 108/236

Suivant