160
4 Molecular Evolution
it represents the true phylogenetic relationships
between the genes and the corresponding species.
Of critical importance to the discussion of the
molecular clock hypothesis is the reliability of the
conclusions about the evolutionary distance
between ancestral forms and their successors that
can be drawn from the lengths of tree connections. However, because the real course of the
phylogenetic history is unknown, the usefulness
of the methods used clearly cannot be judged
from the actual relationships. There is also no justification for any assumption that particular characters of the living organism are a priori superior
for the construction of phylogenetic history such
that the trees based on them could be used as a
standard for the evaluation of other methods.
Consequently, there is no reason at the outset to
assign trees constructed by classical methods,
using morphological characters, any special
degree of correctness. Each phylogenetic tree
represents a hypothesis about the course of evolution, and the plausibility of this hypothesis must
be judged against alternative hypotheses. The
best possible results from attempts to reconstruct
the history of life on earth can actually only be
expected when all available data have been evaluated, be they palaeontological, biogeographical,
morphological or molecular [318].
One way out of the dilemma that the real
course of evolution is unknown and cannot be
used to judge the accuracy of the various reconstruction methods is offered by the application of
the above methods to sets of data obtained in precisely known ways from simulated evolution.
Against such a procedure it can be argued that
simulation of evolution must involve certain basic
assumptions which possibly deviate from the laws
operating in nature. Investigations of simulated
sequences have so far shown that, depending
upon the character of the data, different methods
produce the best results, but also that these
results, even under favourable conditions, may
often be a false branching scheme and involve
large errors in distances [112, 358, 390]. On the
basis of such investigations, particularly heavy
criticism has been levelled against Goodman's
parsimony method [180, 190].
In the ideal case, the same phylogenetic tree
should always emerge for any given species, irrespective of the method used. In actual fact, quite
different trees are often produced when different
methods are applied to the same data or the same
method is used for different sets of molecular
data from the same species [326]. With many
methods, several phylogenetic trees of the same
length (parsimony) are produced from exactly the
same set of data; the introduction of more modern forms changes the branching scheme for the
previously considered species, and the family
relationships between closely related forms
become particularly uncertain. There are various
reasons for these deficiencies [153]:
1. The information contained in the available
data is insufficient. For example, at least 22
characters are required to define a dichotomously branched family tree of 25 species, and
singular (only available for one species) or
incompatible characters are not appropriate.
2. A lack of clarity, caused by polymorphism,
particularly affects the determination of genetic distance between closely related species
[314] .
3. The evolutionary rate plays a significant role.
Where it is low, there are usually too few
evaluable differences; where it is too high,
evolutionary events may be superimposed and
therefore not distinguishable. A variable rate
of evolution complicates the mathematical
process.
4. The largest problem is that of parallel substitution which, contrary to previous notions,
seems to be very frequent. Tests using the
method shown in Fig.4.13 have shown that
sets of molecular data may contain up to 50 %
parallel substitutions [153]. This only applies,
however, to the evaluation of the individual
positions in a sequence; the fear that "almost
or completely identical sequences" may be
found in distantly related species [350] will
probably not be realized.
4.6 The Rate of Molecular Evolution
4.6.1 The Rate of Protein Evolution
The rate of evolution of proteins can be expressed
in several ways: (1) as the rate of amino acid
exchange per amino acid per year; (2) as the time
required for the development of 1 % sequence
difference between two evolving lines (unit
evolutionary period, UEP); and (3) as "accepted
point mutations" (PAM) per 100 amino acids per
100 million years [87, 115, 210,212]. Individual
proteins evolve at very different rates (Table 4.12), with the extremes differing by more than
two orders of magnitude. The neutral theory
explains these differences, in that only certain
4 Molecular Evolution
it represents the true phylogenetic relationships
between the genes and the corresponding species.
Of critical importance to the discussion of the
molecular clock hypothesis is the reliability of the
conclusions about the evolutionary distance
between ancestral forms and their successors that
can be drawn from the lengths of tree connections. However, because the real course of the
phylogenetic history is unknown, the usefulness
of the methods used clearly cannot be judged
from the actual relationships. There is also no justification for any assumption that particular characters of the living organism are a priori superior
for the construction of phylogenetic history such
that the trees based on them could be used as a
standard for the evaluation of other methods.
Consequently, there is no reason at the outset to
assign trees constructed by classical methods,
using morphological characters, any special
degree of correctness. Each phylogenetic tree
represents a hypothesis about the course of evolution, and the plausibility of this hypothesis must
be judged against alternative hypotheses. The
best possible results from attempts to reconstruct
the history of life on earth can actually only be
expected when all available data have been evaluated, be they palaeontological, biogeographical,
morphological or molecular [318].
One way out of the dilemma that the real
course of evolution is unknown and cannot be
used to judge the accuracy of the various reconstruction methods is offered by the application of
the above methods to sets of data obtained in precisely known ways from simulated evolution.
Against such a procedure it can be argued that
simulation of evolution must involve certain basic
assumptions which possibly deviate from the laws
operating in nature. Investigations of simulated
sequences have so far shown that, depending
upon the character of the data, different methods
produce the best results, but also that these
results, even under favourable conditions, may
often be a false branching scheme and involve
large errors in distances [112, 358, 390]. On the
basis of such investigations, particularly heavy
criticism has been levelled against Goodman's
parsimony method [180, 190].
In the ideal case, the same phylogenetic tree
should always emerge for any given species, irrespective of the method used. In actual fact, quite
different trees are often produced when different
methods are applied to the same data or the same
method is used for different sets of molecular
data from the same species [326]. With many
methods, several phylogenetic trees of the same
length (parsimony) are produced from exactly the
same set of data; the introduction of more modern forms changes the branching scheme for the
previously considered species, and the family
relationships between closely related forms
become particularly uncertain. There are various
reasons for these deficiencies [153]:
1. The information contained in the available
data is insufficient. For example, at least 22
characters are required to define a dichotomously branched family tree of 25 species, and
singular (only available for one species) or
incompatible characters are not appropriate.
2. A lack of clarity, caused by polymorphism,
particularly affects the determination of genetic distance between closely related species
[314] .
3. The evolutionary rate plays a significant role.
Where it is low, there are usually too few
evaluable differences; where it is too high,
evolutionary events may be superimposed and
therefore not distinguishable. A variable rate
of evolution complicates the mathematical
process.
4. The largest problem is that of parallel substitution which, contrary to previous notions,
seems to be very frequent. Tests using the
method shown in Fig.4.13 have shown that
sets of molecular data may contain up to 50 %
parallel substitutions [153]. This only applies,
however, to the evaluation of the individual
positions in a sequence; the fear that "almost
or completely identical sequences" may be
found in distantly related species [350] will
probably not be realized.
4.6 The Rate of Molecular Evolution
4.6.1 The Rate of Protein Evolution
The rate of evolution of proteins can be expressed
in several ways: (1) as the rate of amino acid
exchange per amino acid per year; (2) as the time
required for the development of 1 % sequence
difference between two evolving lines (unit
evolutionary period, UEP); and (3) as "accepted
point mutations" (PAM) per 100 amino acids per
100 million years [87, 115, 210,212]. Individual
proteins evolve at very different rates (Table 4.12), with the extremes differing by more than
two orders of magnitude. The neutral theory
explains these differences, in that only certain
