4.1 The Determination of Homology Between Protein DNA Sequences
115
counts only identity, other matrices have been
suggested and in these various amino acid pairs
are weighted, for example, according to differences in the coding triplets, physicochemical similarities, or the relative probability of substitutions. Of the various matrices, the most frequently used is the "mutation data matrix"
(MDM-78) of Dayhoff; this assigns each of the
possible 400 amino acid pairings a specific probability based on analyses of closely related proteins (p. 149). Comparison of the various methods
shows that weighted matrices are only of advantage when the agreement of the compared
sequences is relatively low (below 30 % ). Homology determinations with a higher sensitivity may
be obtained by considering the physicochemical
similarities of amino acids and substitution probabilities [11, 113, 150, 349], or by reference to a
consensus sequence containing all the important
characteristics of a protein family [320].
To achieve optimal agreement it is often necessary to insert "gaps" into one or both sequences,
or to shift one sequence relative to the other so
that unpaired stretches ("tails") are formed. In
this way, however, the chances of random agreement increase. Two random sequences with the
same proportions of all 20 amino acids agree, on
average, by 5 %; deviations in the composition
increase this value to about 6 % for the "average
protein" (see Table 3.1, p.71). If in random
sequences of 100 amino acids relative shifts of up
to 5 amino acids are allowed, the average agreement will have already increased to 8 %; the
introduction of one chosen gap can also increase
the agreement from 5 to 8.5 % [87,93]. Tails and
gaps arise in evolution through the deletion or
insertion of one or more triplet co dons and must
be looked upon as real situations in aligning
sequences. However, their influence on the statistical significance of sequence agreement must
be taken into account; a particular value (a gap
penalty) is usually subtracted from the similarity
factor for each inserted gap [93]. The comparison
of several sequences by alignment in pairs usually
requires various different gaps; this avoids the
method of multiple alignment based on the principle "once a gap, always a gap" [114].
The arithmetical process (algorithm) of
Needleman and Wunsch [286] is most often used
to find the optimal alignment of two sequences.
Orientation of the two sequences at right angles
to each other produces a field on which the similarity between an amino acid of one sequence and
each of the amino acids of the opposing sequence
can be plotted. The pathway across this field that
gives the highest total value of agreement after
subtraction of the gap penalties gives the optimal
alignment (Fig. 4.2). Computer programs are
now used for this manipulation and for the relevant statistical tests. There are also programs for
determining internal periodicities of proteins,
which can arise by duplication of gene segments,
and for searching for related sequences or partial
sequences in databanks. Periodicity is detected by
comparing a partial sequence of a particular
length with all possible partial sequences of that
length in the same protein; in the search for
related sequences, corresponding comparisons
with all partial sequences in a databank are carried out. A further method for aligning sequences
is that of the "dot matrix", in which the "dot criterion" is the number of agreements m in n consecutive amino acids [218]. Due to the continuously
increasing amount of data, such sequence comparisons require increasing mathematical effort.
The development of new or modified computer
programs for molecular research into evolution is
concerned, above all, with simplifying the arithmetic (algorithms) and saving time [29, 96]. It
should then be possible, for example, to compare
{j
I
a
..
V L S P A D K T N V K A A W G K V
V
+
+
H
\
L
T
"'.~
+
P
E
E
K
+~
+
S
A
V +
T
A
L
+
W
G
K
V +
+
+ +
+
+
""-+ +
+
~" +"-+.
+
+
a V-LSPADKTNVKAAWGKV
I I I I I I II II
{j VHLTPEEKSAVTALWGKV
"+
Fig. 4.2. The establishment of the optimal alignment of
two sequences by the method of Needleman and Wunsch
[286], illustrated here for the N-terminal sequences of the
human a- and ~-globin chains. A route that includes the
most + points is sought from the beginning (upper left) to
the end (lower right) of the sequences. Gaps in one or
other sequence are seen as deviations from the diagonal.
The optimal alignment of these two sequences requires a
gap between positions 1 and 2 of the a-chain
115
counts only identity, other matrices have been
suggested and in these various amino acid pairs
are weighted, for example, according to differences in the coding triplets, physicochemical similarities, or the relative probability of substitutions. Of the various matrices, the most frequently used is the "mutation data matrix"
(MDM-78) of Dayhoff; this assigns each of the
possible 400 amino acid pairings a specific probability based on analyses of closely related proteins (p. 149). Comparison of the various methods
shows that weighted matrices are only of advantage when the agreement of the compared
sequences is relatively low (below 30 % ). Homology determinations with a higher sensitivity may
be obtained by considering the physicochemical
similarities of amino acids and substitution probabilities [11, 113, 150, 349], or by reference to a
consensus sequence containing all the important
characteristics of a protein family [320].
To achieve optimal agreement it is often necessary to insert "gaps" into one or both sequences,
or to shift one sequence relative to the other so
that unpaired stretches ("tails") are formed. In
this way, however, the chances of random agreement increase. Two random sequences with the
same proportions of all 20 amino acids agree, on
average, by 5 %; deviations in the composition
increase this value to about 6 % for the "average
protein" (see Table 3.1, p.71). If in random
sequences of 100 amino acids relative shifts of up
to 5 amino acids are allowed, the average agreement will have already increased to 8 %; the
introduction of one chosen gap can also increase
the agreement from 5 to 8.5 % [87,93]. Tails and
gaps arise in evolution through the deletion or
insertion of one or more triplet co dons and must
be looked upon as real situations in aligning
sequences. However, their influence on the statistical significance of sequence agreement must
be taken into account; a particular value (a gap
penalty) is usually subtracted from the similarity
factor for each inserted gap [93]. The comparison
of several sequences by alignment in pairs usually
requires various different gaps; this avoids the
method of multiple alignment based on the principle "once a gap, always a gap" [114].
The arithmetical process (algorithm) of
Needleman and Wunsch [286] is most often used
to find the optimal alignment of two sequences.
Orientation of the two sequences at right angles
to each other produces a field on which the similarity between an amino acid of one sequence and
each of the amino acids of the opposing sequence
can be plotted. The pathway across this field that
gives the highest total value of agreement after
subtraction of the gap penalties gives the optimal
alignment (Fig. 4.2). Computer programs are
now used for this manipulation and for the relevant statistical tests. There are also programs for
determining internal periodicities of proteins,
which can arise by duplication of gene segments,
and for searching for related sequences or partial
sequences in databanks. Periodicity is detected by
comparing a partial sequence of a particular
length with all possible partial sequences of that
length in the same protein; in the search for
related sequences, corresponding comparisons
with all partial sequences in a databank are carried out. A further method for aligning sequences
is that of the "dot matrix", in which the "dot criterion" is the number of agreements m in n consecutive amino acids [218]. Due to the continuously
increasing amount of data, such sequence comparisons require increasing mathematical effort.
The development of new or modified computer
programs for molecular research into evolution is
concerned, above all, with simplifying the arithmetic (algorithms) and saving time [29, 96]. It
should then be possible, for example, to compare
{j
I
a
..
V L S P A D K T N V K A A W G K V
V
+
+
H
\
L
T
"'.~
+
P
E
E
K
+~
+
S
A
V +
T
A
L
+
W
G
K
V +
+
+ +
+
+
""-+ +
+
~" +"-+.
+
+
a V-LSPADKTNVKAAWGKV
I I I I I I II II
{j VHLTPEEKSAVTALWGKV
"+
Fig. 4.2. The establishment of the optimal alignment of
two sequences by the method of Needleman and Wunsch
[286], illustrated here for the N-terminal sequences of the
human a- and ~-globin chains. A route that includes the
most + points is sought from the beginning (upper left) to
the end (lower right) of the sequences. Gaps in one or
other sequence are seen as deviations from the diagonal.
The optimal alignment of these two sequences requires a
gap between positions 1 and 2 of the a-chain
