2.4.1 The DNA Content of Haploid Genomes (the C Value)
25
ing to a divergence of 0.85 % per million years
[283]. The L1 of man, mouse and rats differ in
about 33 % of positions [408]. The conclusion
that there is only one LINE family per mammalian species has become open to question since
the discovery in man of the "L2H" LINE that
deviates greatly from L1H and is found only at
comparable frequencies in the gorilla; in the
chimpanzee it occurs in at least 100 times fewer
copies [315]. The human genome contains at least
40000 L1 copies, and the numbers in other mammals are of the same order. The L1 elements are
not uniformly distributed in the genome: in the
approximately 60 kb of the ~-globin cluster there
are no less than nine L1, whereas with even distribution each L1 would be separated by about
150 kb [399, 470].
2.3.6 Retropseudogenes
The retropseudogenes or processed pseudo genes
are derived almost exclusively from completely
processed RNA and, hence, contain no introns
and have a 3' poly(A) tail. They are very frequent
in mammals; the list of pseudo genes in man, rats
and the mouse includes more than 30 multi-gene
families, e.g. arginine succinate synthase, metallothionein, ~-tubulin, ~-actin, various immunoglobulin chains and oncogenes in humans, as well
as cytochrome c, various ribosomal proteins, atubulin and a-globin in rodents [464, 470]. Like
the other non-viral retroposons, the retropseudogenes are limited mainly to mammals and are
almost unknown in other vertebrates or invertebrates. This is clear from the example of
glyceraldehyde-3-phosphate dehydrogenase, for
which in humans, rabbits, guinea-pigs and hamsters 25 pseudo genes are presently known; in rats
and mice there are actually more than 200, but in
the chicken there is not a single one [470]. Retropseudo genes are transposed passively, which
means that the enzymes for reverse transcription
and insertion must be available in the cell.
Because, under these circumstances, each active
gene could constantly produce additional DNA,
the question arises as to what limits the number
of processed pseudo genes ; it can be assumed that
an equilibrium is established between their
formation and their spontaneous deletion [464].
An important difference exists here between Pol
II- and Pol III-transcribed genes: because polymerase II uses transcription signals outside of the
transcribed region, these are missing in the DNA
copy of the RNA; pseudo genes derived from
mRNA can therefore not be transcribed. In contrast, all the necessary transcription signals of
most Pol-III transcribed genes lie within the transcription unit. Thus, DNA copies of Pol III genes
contain all the information for the production of
further processed pseudogenes [464]. In rate
cases, pseudogenes arise from transcription of an
initially incomplete or aberrantly processed
RNA. If the transcription signals lying in front of
the coding region are also included, then a functional gene copy can result. The only sure example of this at present is the gene for prepro-insulin
in rats and mice. These rodents possess two insulin genes that are more or less equally strongly
expressed. Gene I contains only one intron in the
5' flanking non-coding region, whereas gene II
contains an additional large intron in the
sequence coding for the C peptide. Because the
prepro-insulin gene of the chicken and various
mammals also contains both introns, gene II may
be looked upon as the original rodent gene; gene
I apparently arose from a transcript initiated
approximately 0.5 kb before the normal transcription start site [470]. A functional gene may
also come about if the gene copy produced from
processed mRNA is inserted close behind a Pol II
promotor or later acquires such a promotor. The
intron-free but functional calmodulin gene of the
chicken probably arose in this way [167].
2.4 Size of the Genome
2.4.1 The DNA Content of Haploid Genomes
(the C Value)
Since the first systematic investigations at the
beginning of the 1950s, the DNA content in pg
per haploid genome, the so-called C value, has
been determined for more than 1000 species of
prokaryotes and eukaryotes [62] (Table 2.2). The
DNA content appears to increase with increasing
level of organization. The lowest values in the
animal kingdom are to be found in the Protozoa,
e.g. Plasmodium falciparum, whose genome is
only four to seven times larger than that of E. coli
[168]; extremely large genomes are found
amongst the Urodela. A plausible mechanism for
the increase in genome size during evolution is
genome doubling (polyploidization), followed by
renewed "diploidization" and the diversification
of gene sequences, new arrangements of chromosomes, and loss of parts of the DNA; according to
this process, all eukaryotes would be diploidized
25
ing to a divergence of 0.85 % per million years
[283]. The L1 of man, mouse and rats differ in
about 33 % of positions [408]. The conclusion
that there is only one LINE family per mammalian species has become open to question since
the discovery in man of the "L2H" LINE that
deviates greatly from L1H and is found only at
comparable frequencies in the gorilla; in the
chimpanzee it occurs in at least 100 times fewer
copies [315]. The human genome contains at least
40000 L1 copies, and the numbers in other mammals are of the same order. The L1 elements are
not uniformly distributed in the genome: in the
approximately 60 kb of the ~-globin cluster there
are no less than nine L1, whereas with even distribution each L1 would be separated by about
150 kb [399, 470].
2.3.6 Retropseudogenes
The retropseudogenes or processed pseudo genes
are derived almost exclusively from completely
processed RNA and, hence, contain no introns
and have a 3' poly(A) tail. They are very frequent
in mammals; the list of pseudo genes in man, rats
and the mouse includes more than 30 multi-gene
families, e.g. arginine succinate synthase, metallothionein, ~-tubulin, ~-actin, various immunoglobulin chains and oncogenes in humans, as well
as cytochrome c, various ribosomal proteins, atubulin and a-globin in rodents [464, 470]. Like
the other non-viral retroposons, the retropseudogenes are limited mainly to mammals and are
almost unknown in other vertebrates or invertebrates. This is clear from the example of
glyceraldehyde-3-phosphate dehydrogenase, for
which in humans, rabbits, guinea-pigs and hamsters 25 pseudo genes are presently known; in rats
and mice there are actually more than 200, but in
the chicken there is not a single one [470]. Retropseudo genes are transposed passively, which
means that the enzymes for reverse transcription
and insertion must be available in the cell.
Because, under these circumstances, each active
gene could constantly produce additional DNA,
the question arises as to what limits the number
of processed pseudo genes ; it can be assumed that
an equilibrium is established between their
formation and their spontaneous deletion [464].
An important difference exists here between Pol
II- and Pol III-transcribed genes: because polymerase II uses transcription signals outside of the
transcribed region, these are missing in the DNA
copy of the RNA; pseudo genes derived from
mRNA can therefore not be transcribed. In contrast, all the necessary transcription signals of
most Pol-III transcribed genes lie within the transcription unit. Thus, DNA copies of Pol III genes
contain all the information for the production of
further processed pseudogenes [464]. In rate
cases, pseudogenes arise from transcription of an
initially incomplete or aberrantly processed
RNA. If the transcription signals lying in front of
the coding region are also included, then a functional gene copy can result. The only sure example of this at present is the gene for prepro-insulin
in rats and mice. These rodents possess two insulin genes that are more or less equally strongly
expressed. Gene I contains only one intron in the
5' flanking non-coding region, whereas gene II
contains an additional large intron in the
sequence coding for the C peptide. Because the
prepro-insulin gene of the chicken and various
mammals also contains both introns, gene II may
be looked upon as the original rodent gene; gene
I apparently arose from a transcript initiated
approximately 0.5 kb before the normal transcription start site [470]. A functional gene may
also come about if the gene copy produced from
processed mRNA is inserted close behind a Pol II
promotor or later acquires such a promotor. The
intron-free but functional calmodulin gene of the
chicken probably arose in this way [167].
2.4 Size of the Genome
2.4.1 The DNA Content of Haploid Genomes
(the C Value)
Since the first systematic investigations at the
beginning of the 1950s, the DNA content in pg
per haploid genome, the so-called C value, has
been determined for more than 1000 species of
prokaryotes and eukaryotes [62] (Table 2.2). The
DNA content appears to increase with increasing
level of organization. The lowest values in the
animal kingdom are to be found in the Protozoa,
e.g. Plasmodium falciparum, whose genome is
only four to seven times larger than that of E. coli
[168]; extremely large genomes are found
amongst the Urodela. A plausible mechanism for
the increase in genome size during evolution is
genome doubling (polyploidization), followed by
renewed "diploidization" and the diversification
of gene sequences, new arrangements of chromosomes, and loss of parts of the DNA; according to
this process, all eukaryotes would be diploidized
