6 Identity by Descent in the Mapping of Genetic Traits
143
that we are not here attempting to estimate the pedigree kinship. The histogram of
differences between the pedigree value and actual realized values has a larger spread
than even the upper left histogram based on Equation (6.8).
There is however a more serious deficiency in estimators of the form (6.9); this
is that they do not take the physical locations of SNPs into account. We have
already seen that IBD occurs even in remote relatives as a few long segments.
Additionally, SNPs are individually very uninformative; information about IBD
should be combined across local SNPs to provide more accurate estimates of the
probability of IBD at each point in the genome. Such estimates can then be averaged
across the genome to estimate the realized proportion of genome shared IBD. One
such method was proposed by Day-Williams et al. (2011) and is denoted DW. They
use the four between-individual comparisons of allelic sharing at loci in windows
across the genome, to obtain estimates of local kinship. These are then smoothed
across the genome, subject to constraints that at each point the value is 0, 1/4, 1/2,
or 1 (Table 6.2). An alternative is to estimate the IBD state for the four gametes
using an HMM approach and the model of Brown et al. (2012) for changes in the
15 states across the genome (Sect. 6.2.2): details are given in Sect. 6.3.4 below. This
method provides estimates of realized kinship at points across the genome, and these
may then be combined into a genome-wide estimate.
The two lower histograms of Fig. 6.6 show the results using the local IBD
estimation methods of Day-Williams et al. (2011), denoted DW, and of Brown et al.
(2012) denoted HMM. It is seen that incorporating the segmental nature of DNA
into the inference process greatly improves the precision of estimation of realized
kinship. However, these methods are more computationally intensive and also show
bias. The DW method tends to underestimate IBD, especially in the presence of
inbreeding. The HMM method tends to overestimate IBD in the presence of LD.
For this reason, an LD-weighted version of the HMM estimator provides further
improvement. These and other estimators are further discussed by Wang et al.
(2017).
6.3.4 IBD Given Marker Data in Populations
In the previous section, the focus was on estimating the genome-wide proportions
shared IBD between two individuals. However, for gene mapping, we may be
interested in the joint pattern of IBD among several observed individuals, not
only pairwise measures. Second, for mapping we are interested in IBD at specific
locations across the genome. Third, we may wish to consider segments of IBD
and changes in IBD across genome locations. Estimates of relatedness such as
Equation (6.9) do not take into account the physical linkage among loci, treating
them as an exchangeable collection of SNPs; any permutation of the SNPs will
provide the same result. By contrast, estimates of location-specific IBD rely on the
genetic marker map and depend jointly on the SNPs in the genome region. Each
SNP alone provides little evidence, but segments of IBD typically encompass many
SNPs.
143
that we are not here attempting to estimate the pedigree kinship. The histogram of
differences between the pedigree value and actual realized values has a larger spread
than even the upper left histogram based on Equation (6.8).
There is however a more serious deficiency in estimators of the form (6.9); this
is that they do not take the physical locations of SNPs into account. We have
already seen that IBD occurs even in remote relatives as a few long segments.
Additionally, SNPs are individually very uninformative; information about IBD
should be combined across local SNPs to provide more accurate estimates of the
probability of IBD at each point in the genome. Such estimates can then be averaged
across the genome to estimate the realized proportion of genome shared IBD. One
such method was proposed by Day-Williams et al. (2011) and is denoted DW. They
use the four between-individual comparisons of allelic sharing at loci in windows
across the genome, to obtain estimates of local kinship. These are then smoothed
across the genome, subject to constraints that at each point the value is 0, 1/4, 1/2,
or 1 (Table 6.2). An alternative is to estimate the IBD state for the four gametes
using an HMM approach and the model of Brown et al. (2012) for changes in the
15 states across the genome (Sect. 6.2.2): details are given in Sect. 6.3.4 below. This
method provides estimates of realized kinship at points across the genome, and these
may then be combined into a genome-wide estimate.
The two lower histograms of Fig. 6.6 show the results using the local IBD
estimation methods of Day-Williams et al. (2011), denoted DW, and of Brown et al.
(2012) denoted HMM. It is seen that incorporating the segmental nature of DNA
into the inference process greatly improves the precision of estimation of realized
kinship. However, these methods are more computationally intensive and also show
bias. The DW method tends to underestimate IBD, especially in the presence of
inbreeding. The HMM method tends to overestimate IBD in the presence of LD.
For this reason, an LD-weighted version of the HMM estimator provides further
improvement. These and other estimators are further discussed by Wang et al.
(2017).
6.3.4 IBD Given Marker Data in Populations
In the previous section, the focus was on estimating the genome-wide proportions
shared IBD between two individuals. However, for gene mapping, we may be
interested in the joint pattern of IBD among several observed individuals, not
only pairwise measures. Second, for mapping we are interested in IBD at specific
locations across the genome. Third, we may wish to consider segments of IBD
and changes in IBD across genome locations. Estimates of relatedness such as
Equation (6.9) do not take into account the physical linkage among loci, treating
them as an exchangeable collection of SNPs; any permutation of the SNPs will
provide the same result. By contrast, estimates of location-specific IBD rely on the
genetic marker map and depend jointly on the SNPs in the genome region. Each
SNP alone provides little evidence, but segments of IBD typically encompass many
SNPs.
