116
4 Molecular Evolution
Table 4.1. Pairs of proteins whose homology was tested using the method described on p. 115 [93]. In each case the length
of the compared sequences are given in ( ); the percentage agreement, the number of gaps needed for optimal alignment
and the quotient A [calculated from Eq. (4.1)], which describes the statistical significance, are indicated. The compared
sequences may be considered homologous when A is larger than 3. To calculate the agreement, identical amino acids are
given the value 1 and Cys/Cys pairs the value 2; the gap penalty (p.1l5) was 2.5; "tails" (p.1l5) were not considered
Compared sequences
Human haemoglobin: Hb-~ (146)/Hb-c'\ (146)
Pig lactate dehydrogenase: LDH-M (333)/LDH-H (331)
Human carboanhydrase: CA-B (260)/Ca-C (259)
Bovine chymotrypsinogen-A (245)/trypsinogen (229)
Human haemoglobin: Hb-~ (146)/Hb-a (141)
Human immunoglobulin: CA. (102)/Cx (104)
Bacterial trypsin Streptomyces griseus (221)/
vertebrate trypsin Mustelus canis (222)
Chicken lysozyme (129)/human a-lactalbumin (123)
Human fibrinogen: ~-chain (461)/y-chain (411)
Snake toxin-cardiotoxin Bungarus (118)/
pig prophospholipase (131)
Chicken ovalbumin (386)/human antithrombin III (423)
Carp parvalbumin (108)/bovine troponin-C (161)
Human Hb-a (141)/human myoglobin (153)
Haemoglobin: Chironomus (152)/Myxine (148)
Human cytochrome c (104)1Euglena cytochrome f (87)
Human apolipoprotein AI (125/245)/Corynebacter
diphtheria toxin fragment (125)
Elephant insulin (51)/pig relaxin (48)
Human follitropin (92)/human thyrotropin-~ (112)
Sheep x-casein (171)/human fibrinogen y-chain (179/411)
Bovine chymotrypsin-A (245)/human haptoglobin-~ (245)
Human fibrinogen (N-terminal):
a-chain (239)/y-chain (239)
more than two sequences simultaneously [6, 248,
398,406]. From time to time, completely new
approaches to the comparison of protein sequences are suggested [23, 445]. With the methods
described, it is possible in certain situations to
recognize two amino acid sequences as homologous, although they might differ in more than 75 %
of positions (Table 4.1). Large multi-domain proteins, of which there are many, present particular
problems in the assessment of homology. The
coding sequences of their elements ("modules")
mostly have their origin in different genes that
were recombined by exon shuffling. The change
in function accompanying this incorporation into
a new protein can lead to drastic changes in
sequence which obscure its origin [321].
DNA sequences can be compared by similar
methods [29]. However, significant agreement is
more difficult to show statistically because the
existence of only four different residues already
leads to an average random-sequence agreement
of 25 %. In fact, there are examples where no
homology was detectable at the gene level for
clearly homologous proteins, e.g. the oncogenes
of the sarcoma viruses from the mouse and
Identity
Gaps
A
(%)
93
0
54.0
75
1
75.0
61
1
56.8
46
6
22.6
44
2
17.0
42
3
13.1
38
8
16.9
38
3
10.7
33
5
31.2
32
5
5.0
28
6
14.3
27
2
6.1
27
1
9.3
26
3
3.7
25
3
2.9
25
1
5.2
24
1
4.1
22
2
2.9
21
5
1.4
19
5
3.2
16
2
5.4
chicken [93]. The assessment of homology of noncoding regions is complicated by the frequent
deletions and insertions that occur in association
with the high incidence of substitution.
Homology determination between DNA or
protein sequences is the methodological prerequisite for all investigations of the course and mechanisms of molecular evolution. Of particular
importance is the demonstration that all the proteins of present-day organisms may be arranged
in several hundred groups of homologous sequences, known as protein super-families, after the
suggestion of Margaret Dayhoff. The Atlas of
Protein Sequence and Structure, which appeared
in 1978 [87], listed 248 such protein superfamilies. Several more have been described since
then (Table 4.2), and protein sequences that do
not fit into any of the known families are continuously being discovered; however it is likely that
the total number of super-families is of the order
of 400, or at most 1000. There are many cases of
homology between prokaryote and eukaryote
proteins, e.g. cytochrome, ferredoxin, glyceraldehyde-3-phosphate dehydrogenase, dihydrofolate reductase, trypsin-like enzymes and triose-
4 Molecular Evolution
Table 4.1. Pairs of proteins whose homology was tested using the method described on p. 115 [93]. In each case the length
of the compared sequences are given in ( ); the percentage agreement, the number of gaps needed for optimal alignment
and the quotient A [calculated from Eq. (4.1)], which describes the statistical significance, are indicated. The compared
sequences may be considered homologous when A is larger than 3. To calculate the agreement, identical amino acids are
given the value 1 and Cys/Cys pairs the value 2; the gap penalty (p.1l5) was 2.5; "tails" (p.1l5) were not considered
Compared sequences
Human haemoglobin: Hb-~ (146)/Hb-c'\ (146)
Pig lactate dehydrogenase: LDH-M (333)/LDH-H (331)
Human carboanhydrase: CA-B (260)/Ca-C (259)
Bovine chymotrypsinogen-A (245)/trypsinogen (229)
Human haemoglobin: Hb-~ (146)/Hb-a (141)
Human immunoglobulin: CA. (102)/Cx (104)
Bacterial trypsin Streptomyces griseus (221)/
vertebrate trypsin Mustelus canis (222)
Chicken lysozyme (129)/human a-lactalbumin (123)
Human fibrinogen: ~-chain (461)/y-chain (411)
Snake toxin-cardiotoxin Bungarus (118)/
pig prophospholipase (131)
Chicken ovalbumin (386)/human antithrombin III (423)
Carp parvalbumin (108)/bovine troponin-C (161)
Human Hb-a (141)/human myoglobin (153)
Haemoglobin: Chironomus (152)/Myxine (148)
Human cytochrome c (104)1Euglena cytochrome f (87)
Human apolipoprotein AI (125/245)/Corynebacter
diphtheria toxin fragment (125)
Elephant insulin (51)/pig relaxin (48)
Human follitropin (92)/human thyrotropin-~ (112)
Sheep x-casein (171)/human fibrinogen y-chain (179/411)
Bovine chymotrypsin-A (245)/human haptoglobin-~ (245)
Human fibrinogen (N-terminal):
a-chain (239)/y-chain (239)
more than two sequences simultaneously [6, 248,
398,406]. From time to time, completely new
approaches to the comparison of protein sequences are suggested [23, 445]. With the methods
described, it is possible in certain situations to
recognize two amino acid sequences as homologous, although they might differ in more than 75 %
of positions (Table 4.1). Large multi-domain proteins, of which there are many, present particular
problems in the assessment of homology. The
coding sequences of their elements ("modules")
mostly have their origin in different genes that
were recombined by exon shuffling. The change
in function accompanying this incorporation into
a new protein can lead to drastic changes in
sequence which obscure its origin [321].
DNA sequences can be compared by similar
methods [29]. However, significant agreement is
more difficult to show statistically because the
existence of only four different residues already
leads to an average random-sequence agreement
of 25 %. In fact, there are examples where no
homology was detectable at the gene level for
clearly homologous proteins, e.g. the oncogenes
of the sarcoma viruses from the mouse and
Identity
Gaps
A
(%)
93
0
54.0
75
1
75.0
61
1
56.8
46
6
22.6
44
2
17.0
42
3
13.1
38
8
16.9
38
3
10.7
33
5
31.2
32
5
5.0
28
6
14.3
27
2
6.1
27
1
9.3
26
3
3.7
25
3
2.9
25
1
5.2
24
1
4.1
22
2
2.9
21
5
1.4
19
5
3.2
16
2
5.4
chicken [93]. The assessment of homology of noncoding regions is complicated by the frequent
deletions and insertions that occur in association
with the high incidence of substitution.
Homology determination between DNA or
protein sequences is the methodological prerequisite for all investigations of the course and mechanisms of molecular evolution. Of particular
importance is the demonstration that all the proteins of present-day organisms may be arranged
in several hundred groups of homologous sequences, known as protein super-families, after the
suggestion of Margaret Dayhoff. The Atlas of
Protein Sequence and Structure, which appeared
in 1978 [87], listed 248 such protein superfamilies. Several more have been described since
then (Table 4.2), and protein sequences that do
not fit into any of the known families are continuously being discovered; however it is likely that
the total number of super-families is of the order
of 400, or at most 1000. There are many cases of
homology between prokaryote and eukaryote
proteins, e.g. cytochrome, ferredoxin, glyceraldehyde-3-phosphate dehydrogenase, dihydrofolate reductase, trypsin-like enzymes and triose-
