144
E. A. Thompson
A defined pedigree provides prior probabilities of IBD among individuals at
a locus and across a chromosome (Sect. 6.2.1). However, population genetics
also provides probabilities of coancestry and IBD for individuals sampled from
a population. In the context of modern highly informative SNP data, the prior
distribution has relatively less weight, and genetic marker data can provide strong
evidence of segments of IBD among individuals not known to be related. Formerly,
where marker data were sparse both in the genome and among individuals, the
highly informative pedigree prior was a necessity for successful inference. With
modern data, it is often unnecessarily constraining. Further, ancestral pedigrees may
be inaccurate and cannot be validated from current genetic data. Except among the
current generations of sampled individuals, where the genetic data may be used to
validate the pedigree, the use of a pedigree prior is often best avoided.
Model-based inference of IBD requires allele frequencies, and the relevant allele
frequencies are those in the reference population relative to which IBD is measured.
Sharing of a rare variant allele among individuals provides strong evidence of IBD,
but on average common allelic variation provides more information. Moreover, for
a rare variant, it is difficult to either quantify the strength of the evidence or to
assess its uncertainty; even the concept of a population allele frequency may be
problematic. In contrast, common SNP variation is ancient and relatively stable.
While each SNP alone provides little information, segments of IBD generally
encompass large numbers of SNPs. Whether or not data are phased, and whether
or not local haplotype frequencies are used in estimation, it is the combination of
data from multiple contiguous SNPs that provides the evidence of IBD.
As described in Sect. 6.2.2, a flexible two-parameter prior model for IBD
between the pair of gametes of an individual was introduced by Leutenegger et al.
(2003). In this case there are just two possible IBD states at each locus; the two
gametes are IBD (Z = 1) or they are not (Z = 0). The assumption of a Markov
process of transitions between the two states again gives an HMM structure that
allows inference of IBD segments. The model is shown schematically in Fig. 6.7.
The latent IBD state consists of alternating segments of IBD (Z = 1) and nonIBD (Z = 0) between the two gametes. In an IBD segment, the allelic types
at marker loci are, with high probability, of the same allelic type. In non-IBD
segments they are of independent allelic types. More precisely, the data model
was given in Table 6.3 in Sect. 6.2.4. At each marker locus, given non-IBD, we
have Hardy-Weinberg genotype probabilities. In the case of IBD, a small “error”
probability ε allows for the possibility that IBD DNA may be, or be recorded, as of
Fig. 6.7 Model for inferring IBD between two gametes
E. A. Thompson
A defined pedigree provides prior probabilities of IBD among individuals at
a locus and across a chromosome (Sect. 6.2.1). However, population genetics
also provides probabilities of coancestry and IBD for individuals sampled from
a population. In the context of modern highly informative SNP data, the prior
distribution has relatively less weight, and genetic marker data can provide strong
evidence of segments of IBD among individuals not known to be related. Formerly,
where marker data were sparse both in the genome and among individuals, the
highly informative pedigree prior was a necessity for successful inference. With
modern data, it is often unnecessarily constraining. Further, ancestral pedigrees may
be inaccurate and cannot be validated from current genetic data. Except among the
current generations of sampled individuals, where the genetic data may be used to
validate the pedigree, the use of a pedigree prior is often best avoided.
Model-based inference of IBD requires allele frequencies, and the relevant allele
frequencies are those in the reference population relative to which IBD is measured.
Sharing of a rare variant allele among individuals provides strong evidence of IBD,
but on average common allelic variation provides more information. Moreover, for
a rare variant, it is difficult to either quantify the strength of the evidence or to
assess its uncertainty; even the concept of a population allele frequency may be
problematic. In contrast, common SNP variation is ancient and relatively stable.
While each SNP alone provides little information, segments of IBD generally
encompass large numbers of SNPs. Whether or not data are phased, and whether
or not local haplotype frequencies are used in estimation, it is the combination of
data from multiple contiguous SNPs that provides the evidence of IBD.
As described in Sect. 6.2.2, a flexible two-parameter prior model for IBD
between the pair of gametes of an individual was introduced by Leutenegger et al.
(2003). In this case there are just two possible IBD states at each locus; the two
gametes are IBD (Z = 1) or they are not (Z = 0). The assumption of a Markov
process of transitions between the two states again gives an HMM structure that
allows inference of IBD segments. The model is shown schematically in Fig. 6.7.
The latent IBD state consists of alternating segments of IBD (Z = 1) and nonIBD (Z = 0) between the two gametes. In an IBD segment, the allelic types
at marker loci are, with high probability, of the same allelic type. In non-IBD
segments they are of independent allelic types. More precisely, the data model
was given in Table 6.3 in Sect. 6.2.4. At each marker locus, given non-IBD, we
have Hardy-Weinberg genotype probabilities. In the case of IBD, a small “error”
probability ε allows for the possibility that IBD DNA may be, or be recorded, as of
Fig. 6.7 Model for inferring IBD between two gametes
