56
P. Sjödin et al.
population (given that a pair consists of a female and a male). A consequence
of this is that the probability of an allele contributed by the mother being of
type A is equal to the population frequency of allele type A. The same is true
for alleles contributed by the father. Given a population frequency p of allele
A and 1−p of alleles of type a, the probability of an offspring being of type
AA, Aa, and aa are p 2 , 2p(1−p), and (1−p) 2 , respectively. When this situation
is true, the population is said to be in Hardy-Weinberg equilibrium (HWE).
STRUCTURE has been very popular for population structure inference. However, with the ever-increasing density of genome-wide markers, meeting the computational demands of the algorithm has become a challenge. The Markov chain
Monte Carlo method that STRUCTURE employs places a high burden on computer
resources for large datasets. This has led to the recent development of alternative
approaches, using fast maximum-likelihood-based estimations (FRAPPE (Tang et
al. 2005) and ADMIXTURE (Alexander et al. 2009)).
HAPMIX (Price et al. 2009) extends the local ancestry method implemented in
the second version of STRUCTURE. It is based on the Li and Stephens (2003)
model for patterns of linkage disequilibrium (Li and Stephens 2003) between
markers and infers local ancestry estimates of unphased admixed individuals
based on the phased haplotype data of exactly two populations. Modeling the full
demographic process with recombination and mutation is a notoriously difficult and
computationally intractable problem. The Li and Stephens (2003) approach models
the k + 1 haplotype by (imperfectly) copying from the first k “parental” haplotypes
where recombination events correspond to changing parental haplotype from which
to copy from. Correlations of genealogies (due to linkage) across the sequence is
surprisingly well captured by this approach, and importantly, it is sufficiently simple
to permit even full genome analyses.
Recent developments along this line have led to a method (CHROMOPAINTER,
Lawson et al. (2011)) that does not need discretely defined admixed and parental
populations. Instead, each individual in a sample is considered, in turn, both
as a recipient and a donor, and chromosomes are reconstructed using blocks of
DNA donated by the individuals to each other. Each individual’s chromosome is
thus “painted” by markers donated by donor individuals in any number of other
populations or within the same population. These “painted” chromosomes can be
summarized as a co-ancestry matrix, which is proposed to fully capture the information provided by PCA and STRUCTURE-like methods (also for nonindependent
sites). In addition, consecutive markers that are in linkage disequilibrium are combined into haplotypes, which increases the ability of the method to observe subtle
population structure. A downstream model-based extension (fineSTRUCTURE,
Lawson et al. (2011)) is then used to identify discrete populations using the inferred
co-ancestry matrix.
Précédent

- 62/236

Suivant