50
P. Sjödin et al.
are different. The proportion is obtained by dividing the number of nucleotide
differences by the total number of nucleotides compared. It does not make any
correction for multiple substitutions at the same site, substitution rate biases (for
example, differences in the transition and transversion rates), or differences in
evolutionary rates among sites. More specialized measures (e.g., Jukes-Cantor,
Kimura 2-parameter, Tamura 3-parameter, Tamura-Nei) take some/more of these
complexities into account. The pairwise distances between diploid individuals are
then obtained by averaging over the distances obtained for all pairwise comparisons
of sequences between the two individuals.
There are several tree construction methods for distance data, two of the
most common methods are unweighted pair group method with arithmetic mean
(UPGMA) and neighbor joining (NJ) (Saitou and Nei 1987). UPGMA is a simple
hierarchical clustering method that combines the nearest two units (individuals or
grouped individuals) in a distance matrix into a higher-level cluster. The distance
between any two units is the average of all distances between each element of
each unit. UPGMA assumes a constant molecular clock model and produces an
ultrametric tree (a tree where all the path lengths from the root to the tips are of
equal length). Similar to UPGMA, neighbor joining is also a bottom-up clustering
method; however, compared with UPGMA, neighbor joining has the advantage that
it does not assume that all lineages evolve at the same rate. Both NJ and UPGMA
are fast-clustering (tree-building) algorithms, but since only two elements of the
distance matrix are considered at a time, they have no optimization criterion to fit
the best solution (or tree) over all the data. An optimal criterion method that is
commonly applied to distance data is minimum evolution (ME), which accepts the
tree with the shortest sum of branch lengths, and thus minimizes the total amount of
evolution assumed. Tree-building methods applicable to discrete characters such as
nucleotides are also available (e.g., maximum parsimony and maximum likelihood),
but these are more applicable to phylogenetic purposes and fall outside the scope
of this chapter. Note that inferred trees based on non-recombining chromosomes,
such as the mitochondrial genome or the Y chromosome, represent estimates of the
genealogy of a specific chromosome (see Chap. 1 for a review of gene genealogies)
that may poorly capture an individual’s or a population’s evolutionary history or
structure. Inferred trees based on genome-wide data represent averages over the
genealogical process across the genome, and such summary trees capture population
structure in a more accurate way.
After an initial tree is produced from the distance matrix, a confidence measure
can be calculated making use of procedures such as jackknifing or bootstrapping.
The most widely used tool for confidence inference is a version of bootstrapping
introduced by Felsenstein (1983). Each bootstrap sample consists of the same
number of markers resampled (with replacement) from the original data set and is
then subjected to the same distance calculation and tree reconstruction. From these
trees produced by bootstrapping, a consensus tree can be constructed in which the
confidence of the tree is noted on the nodes as a bootstrap value (the percentage of
times the bootstrap procedure supported the specific node). Jackknife is a similar
resampling procedure, but in this case, the estimate is systematically recomputed by
leaving out one or more observations at a time from the sample set. The bootstrap
Précédent

- 56/236

Suivant