6 Identity by Descent in the Mapping of Genetic Traits
145
different allelic types. While the exact form of the error model is not important, it is
important to allow this flexibility: generally, zero probabilities of latent states should
be avoided in modeling data. Since the model has an HMM structure, standard
algorithms give probabilities Pr(Z j | Y), of the IBD state Z j at each locus j , given
the data X of the allele types on the gametes over all loci. Further, as in Sect. 6.3.2,
we may instead obtain realizations of IBD {Z j ; j = 1, . . . , ,} given X, jointly
over j .
There have been a number of extensions of the basic model, including to analyses
of genotypes of pairs of individuals as implemented in the well-known PLINK
software (Price et al., 2006). A more general prior model for IBD across the genome
and among IBD among any number of gametes was given in Sect. 6.2.2. This model
was used by Brown et al. (2012) on sets of four gametes, using either haplotypic
or genotypic data and by Zheng et al. (2014) for multiple gametes but assuming
haplotypic data. Moltke et al. (2011) also proposed a model for multiple gametes
and used it to study IBD in a set of five individuals. All these approaches have the
same basic framework and objective. That is, patterns of IBD across a chromosome
among sets of gametes are to be inferred from genetic marker data. The IBD process
is approximated by a Markov process, and the allelic or genotypic data at each locus
depends only on the underlying IBD state, giving rise to an HMM.
There is one significant approximation in these models in that linkage disequilibrium (LD) is ignored. That is, there is no direct dependence of allelic types
between loci. While it is the allelic similarity of gametes across multiple loci that
results in inference of IBD, haplotypic similarities are not modeled directly. Allele
frequencies are incorporated into the model, and for common SNP variation, these
are normally adequately accurately known, but haplotype frequencies are often
less well established. While ignoring LD is a model mis-specification that can
result in false inferences of IBD (Fig. 6.6), over-compensation for LD can lead to
failure to detect IBD (Brown et al., 2012). A model that does include LD in the
inference of IBD is that due to Browning and Browning (2010), implemented in
the BEAGLE package but at the expense of a simplified IBD model. This approach
works very well in large samples from large populations, where IBD levels are low
and haplotype frequencies can be well-estimated.
These methods also all face another issue: as the number of gametes n increases,
the number of possible IBD states at each locus increases very rapidly, being the
number of partitions of n items. For the 12 gametes of 6 individuals, there are more
than 4 million possible states. The example considered by Moltke et al. (2011) was
for just five individuals and a limited gene region. Zheng et al. (2014) considered
860 SNP markers over a region of 10 Mbp and succeeded in realizing joint IBD
among 40 gametes but assumed the availability of haplotypic marker data. Neither
of these approaches is scalable to chromosome-wide inferences of joint IBD among
multiple gametes from genotypic marker data.
Précédent

- 149/236

Suivant