4.5.1 The Evolutionary Distance Between Amino Acid or Nucleotide Sequences
149
f)
2)
3)
4)
A
A
~
* ~
A
B
*/\* */\
A
C
B B
A
B
A
A
A
A
* I ~
B
A
*
1\ IV
A
C
B B
A
B
Fig.4.9. Faulty reconstruction of the evolutionary relationship between different characters (A, B and C); this
can result from back mutation (1 and 4), multiple mutation
(2) or parallel mutation (3). The true process is illustrated
in the upper part of the diagram and the mistaken reconstruction is shown in the lower part. The asterisks each
represent a mutation
effective in the evolution of a structural gene but
also the selection for particular structural/functional properties of the DNA and mRNA. Synonymous nucleotide substitutions can, of course,
only be recognized by investigation of the DNA
or mRNA. The number of multiple and back
mutations can be calculated on the basis of a
mathematical model of evolution, the accuracy of
which remains questionable because it cannot
take into account the effect of selection and other
non-stochastic factors. Conclusions drawn about
genes from protein sequences are decreasing in
importance in the wake of advances in nucleic
acid sequencing. The problem of how to include
multiple and back mutations, however, also
applies to nucleotide sequences. To take account
of multiple and back mutations in estimates of
genetic distance using amino acid sequences,
Zuckerkandl and Pauling suggested, in 1965, the
use of a Poisson correction. Accordingly, the
average number of exchanges per amino acid
between two sequences is given by
Kaa = -loge[l-(daa/naa)] = -loge(l-Pd) (4.10)
and the rate of protein evolution is given by
kaa = Kaal(2T)
(4.11)
In these equations, Pd is the proportion of varying
amino acids, calculated from the number of
amino acid differences (daa) and the length of the
sequences (naa) and Tis the separation time of the
evolution lines of the two sequences. The equation gives usable results for up to approximately
40 % sequence difference (Pd :5 0.40) [210, 212].
For the determination of genetic difference in the
case of larger differences, Margaret Dayhoff
derived a matrix using 1572 amino acid exchanges
in closely related proteins; this gives the relative
probabilities of individual amino acid exchanges
("mutation data matrix" MDM78). With the help
of this matrix it is possible to calculate a genetic
distance for each sequence pair of the amino acid
differences, and for this purpose the distance
measurement PAM ("point accepted mutations" ,
defined as substitutions per 100 co dons ) is introduced [87]. The procedure has been recently
improved [435]. Numerically similar values are
obtained with the following equation [210]:
(4.12)
A new matrix (EMPAR, exchange matrix derived
from parameters) similar to the MDM78 , has now
been suggested and makes use of certain physicochemical parameters of the amino acids, instead
of observations on amino acid exchanges. As
evolution prefers conservative amino acid
exchanges in which the physicochemical properties are only slightly altered, the EMPAR and the
MDM78 matrices assign individual amino acid
exchanges similar probabilities [342]. Since 1972,
Holmquist and Jukes have been working on the
development of a measurement of distance
(REH, "random evolutionary hits") from a stochastic model of evolution; this makes allowance
for all substitutions, including synonymous,
multiple and back substitutions. This model assumes that at any point during evolution only some
of the co dons vary, and these "varions" change
after each substitution [141]. The maximum parsimony method, developed by Goodman and
Moore from 1972 to 1976, starts with a phylogenetic tree of minimal total length ("maximum parsimony tree"). The lower the "density" of any
part of this tree, i.e. the lower the amount of
sequence information to be evaluated, the more
the frequency of multiple and back mutations will
be underestimated. This is taken into account by
the calculation of an "augmented distance" (AD)
[141].
The AD method always makes use of a directly
determined nucleotide sequence, or one derived
from an amino acid sequence. The REH method
in a modified form can also use nucleic acid data
[141]. In addition to the possibility of multiple
Précédent

- 164/799

Suivant