164
B. F. Voight
turn out to be, especially given the number of not only markers directly assayed
and tested but also implicit (but partially correlated) tests of additional markers
not assayed but indirectly tested due to linkage disequilibrium? What emerged
was a consensus of observations from multiple lines of inquiry designed to
estimate the threshold empirically from data—permutations, principal components
methods, haplotype-block counting, and more—all of which triangulated on a
threshold around 5 × 10 −8 . This specific value maintained equal to or slightly
lower than the required error rates given the number and haplotype structure
around common variants tested (Johnson et al. 2010). While discussion about
the appropriateness of this specific choice continues today, the emergence of an
overall threshold allowing specific findings to stand out in the context of the entire
genome (rather than candidate genes) was critical to identify bona fide observations
which could be statistically comparable across studies and across phenotypes,
observations that would stand up to scrutiny and replication efforts. The fact that
now thousands of associations exceed this threshold and have been independently
replicated, across multiple ethnicities (Saxena et al. 2012), validates the central
claim that GWAS can identify common genetic variants that contribute to complex
traits.
7.2.4 Best Practices for In Silico Statistical Imputation
While the direct testing of markers on genotyping arrays was a major advance
to understanding complex disease, two central challenges emerged as multiple
technologies, and data sets emerged to perform the task. First, after the issue of
statistical thresholds had been addressed, direct testing as many of the ∼2.5 million
markers found in population databases (International HapMap Consortium 2005,
2007) would likely be the best shot at detecting association with any SNP. Second, as
new studies increasingly used different genotyping technologies often with different
and only partially overlapping SNP panels, it was clear that a way to summarize
evidence of association across a common SNP panel would be desirable. These two
facts and an emergently clear picture about the landscape of haplotypes in genetic
data motivated the development of statistical methods that use panels of reference
haplotypes from population studies to statistically infer the genotypes of untyped
markers. The intuition for how this process works comes fundamental to the idea
of linkage disequilibrium, in a multi-locus context. If one can predict with high
accuracy (r 2 > 0.95, for example) the genotype of one SNP given another, one does
not need to know the precise SNP genotypes at one site to test the other. Now,
consider that information on linkage extends beyond simply pairs of SNPs, but to
multiple SNPs that exist on specific haplotypes. If one can match the haplotype in
question closely to combinations of those that have been previously observed, then
one can probabilistically infer the genotypes at all untyped markers that also reside
on those haplotypes with relatively high accuracy. This process (called imputation)
now has many robust and specific implementations in software (Browning and
Browning 2009; Howie et al. 2009; Li et al. 2010). The principles of how best to
Précédent

- 168/236

Suivant