134
E. A. Thompson
For one between-individual IBD link (states 11, 12, 13, and 14), the value is 1/4. If
all four gametes are IBD, the value is 1, and the remaining 6 states each gives value
1/2. From Equation (6.4) and its analogues for other gamete pairs, it is seen that,
under the ESF model, β is both the kinship between A and B and the inbreeding
coefficient of each individual (the probability of IBD between the individual’s two
gametes).
The IBD state among gametes changes along a chromosome due to ancestral
recombination events. For a pair of gametes, Leutenegger et al. (2003) proposed a
simple model in which potential changes occur at rate α and at a potential change
point the new (possibly unchanged) state is IBD or non-IBD with probability β and
(1 − β), respectively. This gives rise to an equilibrium pairwise IBD probability β
and to alternating segments of IBD and non-IBD. The lengths of these segments are
exponentially distributed with expectations 1/α(1 − β) and 1/αβ, respectively. This
model is based on the the consideration of a single chain of ancestry (Fig. 6.4) and
does not reflect a more complex situation where there are multiple paths of ancestry
of varying numbers of meioses between the two gametes. However, it is flexible
enough to provide a useful prior distribution for IBD.
Modeling the changes in IBD among multiple gametes along a chromosome is
a challenging problem. The full ancestral recombination graph is too complex a
model for genome-wide use. Simple approximations cannot accommodate the range
of changes that can occur or fail to mimic the types of changes that do occur. For
example, an extension of the model of Leutenegger et al. (2003), which samples
from the ESF at each potential change point, would allow immediate changes from
state 1 to state 15, in Table 6.2, or from state 9 to state 10, whereas no single ancestral
recombination event could accomplish these changes.
One model that has proved useful is that proposed by Brown et al. (2012) which
applies to any number of gametes and has the ESF as its equilibrium distribution.
This model allows for the move of any one gamete into, out of, or between any
two IBD subsets at each potential transition point. Potential transitions occurs at
rate α, which is a surrogate for recombination rate. The two parameters β and α
together control the overall level of IBD and the lengths of chromosome over which
a subset of gametes will remain IBD. This model also does not accommodate all
possible transitions. For example, an ancestral recombination that is ancestral to
two current gametes may move them together into another IBD subset. However,
provided other changes are allowed for with some small probability, this model also
provides a useful prior (Zheng et al., 2014).
6.2.3 Probabilities of Genotypic Data Given IBD
We now consider the relationship between latent IBD and the probabilities of marker
genotypes. At a specific locus, in a specific gamete, the probability of the allelic type
of the DNA is simply the population allele frequency. We will label SNP alleles as
u and v, with population frequencies q and (1 − q). At any locus the observed
genotype G(A) of any individual A is thus uu, uv, or vv. More generally, for a
E. A. Thompson
For one between-individual IBD link (states 11, 12, 13, and 14), the value is 1/4. If
all four gametes are IBD, the value is 1, and the remaining 6 states each gives value
1/2. From Equation (6.4) and its analogues for other gamete pairs, it is seen that,
under the ESF model, β is both the kinship between A and B and the inbreeding
coefficient of each individual (the probability of IBD between the individual’s two
gametes).
The IBD state among gametes changes along a chromosome due to ancestral
recombination events. For a pair of gametes, Leutenegger et al. (2003) proposed a
simple model in which potential changes occur at rate α and at a potential change
point the new (possibly unchanged) state is IBD or non-IBD with probability β and
(1 − β), respectively. This gives rise to an equilibrium pairwise IBD probability β
and to alternating segments of IBD and non-IBD. The lengths of these segments are
exponentially distributed with expectations 1/α(1 − β) and 1/αβ, respectively. This
model is based on the the consideration of a single chain of ancestry (Fig. 6.4) and
does not reflect a more complex situation where there are multiple paths of ancestry
of varying numbers of meioses between the two gametes. However, it is flexible
enough to provide a useful prior distribution for IBD.
Modeling the changes in IBD among multiple gametes along a chromosome is
a challenging problem. The full ancestral recombination graph is too complex a
model for genome-wide use. Simple approximations cannot accommodate the range
of changes that can occur or fail to mimic the types of changes that do occur. For
example, an extension of the model of Leutenegger et al. (2003), which samples
from the ESF at each potential change point, would allow immediate changes from
state 1 to state 15, in Table 6.2, or from state 9 to state 10, whereas no single ancestral
recombination event could accomplish these changes.
One model that has proved useful is that proposed by Brown et al. (2012) which
applies to any number of gametes and has the ESF as its equilibrium distribution.
This model allows for the move of any one gamete into, out of, or between any
two IBD subsets at each potential transition point. Potential transitions occurs at
rate α, which is a surrogate for recombination rate. The two parameters β and α
together control the overall level of IBD and the lengths of chromosome over which
a subset of gametes will remain IBD. This model also does not accommodate all
possible transitions. For example, an ancestral recombination that is ancestral to
two current gametes may move them together into another IBD subset. However,
provided other changes are allowed for with some small probability, this model also
provides a useful prior (Zheng et al., 2014).
6.2.3 Probabilities of Genotypic Data Given IBD
We now consider the relationship between latent IBD and the probabilities of marker
genotypes. At a specific locus, in a specific gamete, the probability of the allelic type
of the DNA is simply the population allele frequency. We will label SNP alleles as
u and v, with population frequencies q and (1 − q). At any locus the observed
genotype G(A) of any individual A is thus uu, uv, or vv. More generally, for a
