156
E. A. Thompson
that are of order of millions of base pairs long and will contain many SNP markers.
Also in Sect. 6.2 we considered probabilities of marker data, X, and trait data, Y,
conditionally on Z. The key point here is that, for these probabilities, it is irrelevant
whether or not the pedigree structure is known.
In Sect. 6.3 we turn to the inference of Z given marker data X. We first consider
the case of defined relatives, where the pedigree structure provides a strong, but
sometimes overly constraining, prior. Usually not all members of a pedigrees are
typed, but pedigrees can only be validated among current individuals for whom
marker genotypes are available. In some studies also there may be issues of
genotypic error; for computational reasons, pedigree-based methods assume that
marker data are observed without error. We then turn to the inference of IBD in
population, the purpose of the IBD model being to provide a flexible and tractable
prior for inference. One important aspect of this flexibility is that is allows for
error in the observation of marker genotypes. SNP typing is quite accurate, but
there are very large numbers of SNPs. We consider first genome-wide measures
of IBD. We point out issues with methods that treat all SNPs equally and take
no account of their dependence due either to allelic association (LD) or to their
physical locations. We describe one method of adjusting for LD, but our major
focus again is on the segmental natural of DNA descent. Since individual SNPs are
very uninformative, combining multiple SNPs in detecting IBD segments is of key
importance. Using a model for the dependence of descent across SNP markers has
two important consequences. First, estimates of genome-wide IBD proportions are
greatly improved. Second, and essential for gene mapping purposes, Z is realized at
locations across the genome: the actual locations of segments of IBD are detected.
Finally, in Sect. 6.4 we show how Z inferred from marker data X can be used to
map DNA that is causal to trait data Y against the genetic marker map. Again we
consider first the case of pairs or groups of individuals whose pedigree relationships
are known. This includes approaches such as that of affected relative pairs for binary
traits and variance component models for mapping quantitative trait loci (QTL). We
then extend to more general models for phenotypes Y and show how realizations
of Z conditional on X can be used to obtain Monte Carlo estimates of linkage
likelihoods Pr(Y|X) for a specified trait model and specified hypotheses of the
location(s) of causal DNA. Next we return to populations and consider an IBDbased analogue of case-control studies, showing that where different rare variants
in a functional gene can cause the trait, the IBD-based approach outperforms a
GWAS test. IBD-based tests can address allelic heterogeneity both in pedigrees
and in populations. Finally, we return to linkage likelihoods, on the basis of IBD
inferred in populations where pedigree relationships are unknown. We consider the
complexities of multi-individual IBD and suggest that often a variance component
model that requires only pairwise IBD may be more useful. However, the more
fundamental message is that it is largely irrelevant to subsequent analysis whether
IBD is inferred under a population model or on a defined pedigree. All pedigrees
exist within a broader population framework; the IBD framework permits the
combination of population and pedigree information.
E. A. Thompson
that are of order of millions of base pairs long and will contain many SNP markers.
Also in Sect. 6.2 we considered probabilities of marker data, X, and trait data, Y,
conditionally on Z. The key point here is that, for these probabilities, it is irrelevant
whether or not the pedigree structure is known.
In Sect. 6.3 we turn to the inference of Z given marker data X. We first consider
the case of defined relatives, where the pedigree structure provides a strong, but
sometimes overly constraining, prior. Usually not all members of a pedigrees are
typed, but pedigrees can only be validated among current individuals for whom
marker genotypes are available. In some studies also there may be issues of
genotypic error; for computational reasons, pedigree-based methods assume that
marker data are observed without error. We then turn to the inference of IBD in
population, the purpose of the IBD model being to provide a flexible and tractable
prior for inference. One important aspect of this flexibility is that is allows for
error in the observation of marker genotypes. SNP typing is quite accurate, but
there are very large numbers of SNPs. We consider first genome-wide measures
of IBD. We point out issues with methods that treat all SNPs equally and take
no account of their dependence due either to allelic association (LD) or to their
physical locations. We describe one method of adjusting for LD, but our major
focus again is on the segmental natural of DNA descent. Since individual SNPs are
very uninformative, combining multiple SNPs in detecting IBD segments is of key
importance. Using a model for the dependence of descent across SNP markers has
two important consequences. First, estimates of genome-wide IBD proportions are
greatly improved. Second, and essential for gene mapping purposes, Z is realized at
locations across the genome: the actual locations of segments of IBD are detected.
Finally, in Sect. 6.4 we show how Z inferred from marker data X can be used to
map DNA that is causal to trait data Y against the genetic marker map. Again we
consider first the case of pairs or groups of individuals whose pedigree relationships
are known. This includes approaches such as that of affected relative pairs for binary
traits and variance component models for mapping quantitative trait loci (QTL). We
then extend to more general models for phenotypes Y and show how realizations
of Z conditional on X can be used to obtain Monte Carlo estimates of linkage
likelihoods Pr(Y|X) for a specified trait model and specified hypotheses of the
location(s) of causal DNA. Next we return to populations and consider an IBDbased analogue of case-control studies, showing that where different rare variants
in a functional gene can cause the trait, the IBD-based approach outperforms a
GWAS test. IBD-based tests can address allelic heterogeneity both in pedigrees
and in populations. Finally, we return to linkage likelihoods, on the basis of IBD
inferred in populations where pedigree relationships are unknown. We consider the
complexities of multi-individual IBD and suggest that often a variance component
model that requires only pairwise IBD may be more useful. However, the more
fundamental message is that it is largely irrelevant to subsequent analysis whether
IBD is inferred under a population model or on a defined pedigree. All pedigrees
exist within a broader population framework; the IBD framework permits the
combination of population and pedigree information.
