exceeds that of functional genes, the proportions
being, for example, 4:1 for the cytochrome c
gene of rats, 8: 3 for the tubulin genes, 10: 1 for
the U1 RNA genes of man, and between 6: 1 and
32: 1 for various ribosomal protein genes of the
mouse [464]. Two categories of pseudo genes are
recognized according to their origins: the "traditional" and the "processed" pseudogenes. The
traditional pseudogenes arose through the
duplication of gene copies that, after a period of
normal function, became defective, e.g. by single
base substitutions that introduce intrinsic stop
codons (nonsense mutations) or by the insertion
or deletion of one or two bases that shift the reading frame (frameshift mutations). Genes may also
be inactivated by changes in transcription or
translation signals or by substitutions in the control region.
The pseudo genes of the second category are
distinguished from functional genes by the
absence of introns and by their remote location in
the genome, far from their functional counterparts. Many possess a 3' terminal poly(A)
sequence of 11-38 bp; many are flanked by identical, similarly oriented sequences (direct repeats).
All these characteristics speak for a mode of origin whereby a DNA copy produced by reverse
transcription of a finished (processed) mRNA is
inserted somewhere in the genome. The processed pseudo genes thus belong to the large number of RNA-dependent transposed sequences
(retroposens), and are also termed retropseudogenes. This type of pseudo gene is found in many
multi-gene families of mammals but, up to now,
has been found only rarely in other vertebrates
and the invertebrates. Only by sequence analysis
will it be possible to decide whether the widely
spread orphons are really pseudogenes of this
category [452,464, 470]. The retropseudogenes
and the other retroposons will be discussed more
thoroughly below.
2.3 Repetitive Sequences
and Mobile Elements
Whereas the small genomes of the prokaryotes
consist almost entirely of unique sequences, the
much larger eukaryote genomes consistently contain a proportion of repetitive sequences, which
in unicellular or multicellular animals may be
between 10 and 50 % of the DNA. The proportion of repetitive, information-deficient sequences in DNA fragments can be estimated from the
2.3 Repetitive Sequences and Mobile Elements
17
rate with which single strands reassociate to
duplexes (reassociation kinetics). The product of
the starting concentration and half-life, cot,
increases with increasing complexity. The cot
value distinguishes highly repetitive DNA with
about 10 6 copies, middle repetitive DNA with
10 2 _10 5 copies, and single-copy (unique) DNA
with one or a few copies. The multi-gene families
of the ribosomal RNAs and histones fall into the
category of middle repetitive DNA, although
only with proportions of a few percent. The Protozoa, with their small genomes, have only
10-20 % repetitive DNA [203]; amongst the
vertebrates, the teleost Arothron diadematus has
the lowest proportion of repetitive DNA with
13 % [345]; in contrast, a relatively large proportion, particularly of highly repetitive DNA, is
found in the extremely large genomes of the Urodela [43].
Many of the highly repeated sequences form
tandem clusters in the heterochromatin of the
centromere and telomere regions of chromosomes (satellite DNA). Tandem repeats of short
sequences are also found in the telomeres themselves, the specialized structures at the ends of
chromosomes that play an important role through
their replication, stabilization and interaction
with the nuclear membrane. The molecular structure of the telomeres was first investigated in the
Ciliophora where 10 4 -10 7 of such chromosome
ends are to be found in the macronucleus. Telomeric repeats of (TTGGGG)m were found in Tetrahymena, and of (TTTTGGGG)n in the hypotrichous Ciliophora Oxytricha, Euplotes and Stylonchia. The telomeres have a single-stranded
extension to which specific proteins are bound
[192]. In addition to the repeats (TTGGGG) and
(TGAGGG), one finds especially (TTAGGG) in
humans, all classes of vertebrates, and the flagellate trypanosomes. The combined length of the
tandem repeats in the lower eukaryotes amounts
at the most to 1 kb, but in man it is 10-15 kb, and
in the mouse it is up to 100 kb [36, 491]. Short,
highly repeated sequences of only 2-3 bp are also
widely scattered in the genome, e.g. in introns, in
the spacers between genes, and also within longer
repeated sequences. All possible types of such
short, dispersed sequences are found in very different classes of eukaryotes (man, Drosophila,
sea urchin, the micronucleus of the ciliate Stylonchia, and yeast); examples include repeats such
as AA!IT (i.e. poly(A) on one strand and
poly(T) on the other), GT/CA or CAG/GTC.
Unlike similar satellite-DNA sequences, these
dispersed sequences are transcribed [432]. Droso-
being, for example, 4:1 for the cytochrome c
gene of rats, 8: 3 for the tubulin genes, 10: 1 for
the U1 RNA genes of man, and between 6: 1 and
32: 1 for various ribosomal protein genes of the
mouse [464]. Two categories of pseudo genes are
recognized according to their origins: the "traditional" and the "processed" pseudogenes. The
traditional pseudogenes arose through the
duplication of gene copies that, after a period of
normal function, became defective, e.g. by single
base substitutions that introduce intrinsic stop
codons (nonsense mutations) or by the insertion
or deletion of one or two bases that shift the reading frame (frameshift mutations). Genes may also
be inactivated by changes in transcription or
translation signals or by substitutions in the control region.
The pseudo genes of the second category are
distinguished from functional genes by the
absence of introns and by their remote location in
the genome, far from their functional counterparts. Many possess a 3' terminal poly(A)
sequence of 11-38 bp; many are flanked by identical, similarly oriented sequences (direct repeats).
All these characteristics speak for a mode of origin whereby a DNA copy produced by reverse
transcription of a finished (processed) mRNA is
inserted somewhere in the genome. The processed pseudo genes thus belong to the large number of RNA-dependent transposed sequences
(retroposens), and are also termed retropseudogenes. This type of pseudo gene is found in many
multi-gene families of mammals but, up to now,
has been found only rarely in other vertebrates
and the invertebrates. Only by sequence analysis
will it be possible to decide whether the widely
spread orphons are really pseudogenes of this
category [452,464, 470]. The retropseudogenes
and the other retroposons will be discussed more
thoroughly below.
2.3 Repetitive Sequences
and Mobile Elements
Whereas the small genomes of the prokaryotes
consist almost entirely of unique sequences, the
much larger eukaryote genomes consistently contain a proportion of repetitive sequences, which
in unicellular or multicellular animals may be
between 10 and 50 % of the DNA. The proportion of repetitive, information-deficient sequences in DNA fragments can be estimated from the
2.3 Repetitive Sequences and Mobile Elements
17
rate with which single strands reassociate to
duplexes (reassociation kinetics). The product of
the starting concentration and half-life, cot,
increases with increasing complexity. The cot
value distinguishes highly repetitive DNA with
about 10 6 copies, middle repetitive DNA with
10 2 _10 5 copies, and single-copy (unique) DNA
with one or a few copies. The multi-gene families
of the ribosomal RNAs and histones fall into the
category of middle repetitive DNA, although
only with proportions of a few percent. The Protozoa, with their small genomes, have only
10-20 % repetitive DNA [203]; amongst the
vertebrates, the teleost Arothron diadematus has
the lowest proportion of repetitive DNA with
13 % [345]; in contrast, a relatively large proportion, particularly of highly repetitive DNA, is
found in the extremely large genomes of the Urodela [43].
Many of the highly repeated sequences form
tandem clusters in the heterochromatin of the
centromere and telomere regions of chromosomes (satellite DNA). Tandem repeats of short
sequences are also found in the telomeres themselves, the specialized structures at the ends of
chromosomes that play an important role through
their replication, stabilization and interaction
with the nuclear membrane. The molecular structure of the telomeres was first investigated in the
Ciliophora where 10 4 -10 7 of such chromosome
ends are to be found in the macronucleus. Telomeric repeats of (TTGGGG)m were found in Tetrahymena, and of (TTTTGGGG)n in the hypotrichous Ciliophora Oxytricha, Euplotes and Stylonchia. The telomeres have a single-stranded
extension to which specific proteins are bound
[192]. In addition to the repeats (TTGGGG) and
(TGAGGG), one finds especially (TTAGGG) in
humans, all classes of vertebrates, and the flagellate trypanosomes. The combined length of the
tandem repeats in the lower eukaryotes amounts
at the most to 1 kb, but in man it is 10-15 kb, and
in the mouse it is up to 100 kb [36, 491]. Short,
highly repeated sequences of only 2-3 bp are also
widely scattered in the genome, e.g. in introns, in
the spacers between genes, and also within longer
repeated sequences. All possible types of such
short, dispersed sequences are found in very different classes of eukaryotes (man, Drosophila,
sea urchin, the micronucleus of the ciliate Stylonchia, and yeast); examples include repeats such
as AA!IT (i.e. poly(A) on one strand and
poly(T) on the other), GT/CA or CAG/GTC.
Unlike similar satellite-DNA sequences, these
dispersed sequences are transcribed [432]. Droso-
