138
E. A. Thompson
SNP genotypes, another model is that each allele is independently toggled to
its alternative with probability ε (Zheng et al., 2014). Since some loci are more
error-prone, ε may be made locus-dependent. With this or a more general error
model, genetic marker become in effect discrete trait phenotypes. Then probability
computations require the general summation method exemplified in Equation (6.5).
6.3
Inferring in Pedigrees and Populations
6.3.1 IBD Given Marker Data on Relatives
In this section we will consider the inference of IBD from genetic marker data.
We consider first the case where the pedigree structure of the observed individuals
is known and assumed correct. Marker genotypes for some individuals for some
subsets of loci may be missing, but we assume that, if observed, the marker
genotypes are without error. As an example of the principles involved, we consider
the example of a pair of full sibs. At any locus, sibs share their maternal/paternal
genome IBD with probability 1/2. In the absence of genetic marker data, there are
prior probabilities 1/4, 1/2, and 1/4 (respectively) that they share 0, 1, or 2 gene
copies IBD (Sect. 6.2.1).
More generally, we will denote genetic marker (usually SNP) data by X and a
specification of the IBD to be inferred by Z. Section 6.2.3 showed how probabilities
Pr(X | Z) of genetic marker data on relatives could be easily computed given the a
pattern of IBD at the marker locus. Conversely, given a prior probability Pr(Z) of
IBD and genetic marker data, conditional probabilities of IBD can be obtained. By
Bayes theorem:
Pr(Z | X) ∝ Pr(X | Z) Pr(Z).
(6.7)
For pairs of individuals in a known pedigree relationship, there are well-established
methods for computing these prior probabilities (Karigl, 1981).
Suppose the SNP genotypes of the pair of sibs at three linked loci are as shown in
Table 6.4. As in Sect. 6.2.3 we will denote the SNP alleles as u and v, while q j and
(1 − q j ) will now denote the frequency of u and of v at locus j . The reader should
not struggle with details of the computation but consider only whether the results
make qualitative sense. Given the allele frequencies shown, then the probabilities
that the sibs share 0, 1, or 2 DNA copies IBD are computed using Equation (6.7)
and are given in Table 6.4. Note the effect of the u allele frequency, q j . The two sibs
have the same genotypes at locus-1 and locus-3, but at locus-3 the u allele is the
rare allele, giving much stronger weight to IBD sharing between the two sibs. Only
at locus-2 do the two sibs have the same genotype and so can share 2 copies IBD
(Z = 2). The genotypic data raises the probability that they do so to 0.4, which is
higher than above the prior probability 0.25.
However, if the loci are linked, this computation does not take into account all the
information available; there is dependence in the IBD state Z at linked loci. Suppose
Précédent

- 142/236

Suivant