found in only some of the multiple 28S rRNA
genes. The insulin genes of the vertebrates consistently have two introns, one in the 5' non-coding
region and one in the region of the C-peptide; in
rats, however, the C-peptide intron is missing
from one of the two insulin genes present [45].
In some introns there are open reading frames
(ORFs) without stop codons that can code for
proteins. The first example of an exon in an
intron was the maturase gene, discovered within
the yeast cytochrome-b gene in 1981; this plays a
role in the splicing of the cytochrome-b transcript. Two examples are known in Drosophila:
The first is the gart gene, which codes for three
enzymes of purine synthesis and contains in its
first 4142-bp-long intron on the complementary
strand, i.e. being read in the opposite direction,
the fully functional gene for pupa cuticular protein (PCP) with 184 amino acids, a homologue of
four known larval cuticular proteins (LCPs). In
contrast to the gart gene, the PCP gene is expressed only in the epidermis of the pre-pupa. The
PCP gene situated in the gart intron itself contains an intron of 71 bp between co dons 4 and 5.
A sequence comparison with the relevant yeast
gene clearly shows that the PCP gene was
inserted later [189]. Sequence comparisons suggest that the PCP gene arose about 70 million
years ago by duplication of the then single LCP
gene and is, therefore, much older than the present Drosophila species [307]. The second example is the molecular organization of the dounce
gene which codes for a cAMP phosphodiesterase
and has proven to be particularly complex, with
three divergently transcribed genes nested within
its introns. Two of these, Sgs-4 and Pig-1, are
nested within the extremely large 79-kb intron
that separates exon 3 from exon 2 [69, 143]. Corresponding examples are also to be found in
mammals and man. A gene of unknown function
is located in the first intron of the mouse (3glucoronidase gene; introns 1 and 3 of the human
pre-albumin gene contain ORFs that could code
for proteins of 37 to 69 amino acids [442, 465].
There are two basically different hypotheses
about intron evolution that relate to the localization of introns within coding sequences and the
splicing process. According to the "association"
hypothesis, introns and splicing were already present in the first organisms but were lost in the prokaryotes when DNA content was reduced in
favour of more rapid replication. The second
hypothesis ("divisive" intron origin) assumes that
coding sequences originally were continuous and
lacking in introns, and that only during eukaryote
2.1.4 Introns
15
evolution were introns introduced and the splicing mechanism developed. Here one can think of
introns arising from transposable elements that
carried on their ends signals specific for insertion
into DNA sequences that were already present
and for the splicing process. The first hypothesis
assumes that longer coding sequences arose during the early evolution of organisms by the joining of shorter sequences. Speaking in favour of
this is the existence in prokaryotic genomes of a
primitive, RNA-dependent splicing mechanism
and the albeit rare occurrence of introns. A particularly strong argument for the associative origin of introns is that in many genes single exons
represent subregions of polypeptide chains that
either fold in a compact manner (modules) or
are structurally and functionally autonomous
(domains). On the other hand, for many genes
there is no correlation between intron position
and domain structure. The appearance of new
introns can be demonstrated in the evolution of
the serine proteases (see Fig. 3.6; p.91). The
tubulin and actin genes of eukaryotes vary
enormously in the number and position of
introns: a- and (3-tubulin are homologous and
agree in about 40 % of their amino acids, but only
one of the 17 possible intron positions in atubulin and the 19 possible intron positions in (3tubulin coincide. All these positions, however,
regardless of whether they contain an intron or
not, show characteristics of the typical intronflanking "junction sequence"; they appear to be
"proto-splice sites" predestined for the insertion
of introns [107, 155,421].
The rearrangement of exons (exon shuffling)
has played a large role in molecular evolution
[366,421]. One finds, for example, in genes of
very different function and origin, one or more
exons that code for an amino acid sequence
homologous to epidermal growth factor (EGF)
(see Fig. 3.6; p. 91). That introns may be very old
is nicely demonstrated by the genes of triosephosphate isomerase; these are homologous in all
organisms. In E. coli and baker's yeast these
genes have no introns, but they have five introns
in the fungus Aspergillus nidulans, six in vertebrates, and eight in maize. Five of the introns have
the same position in maize and man. Many of the
introns lie in regions of the triosephosphate isomerase genes whose amino acid sequence is
highly conserved in all organisms; this speaks
against them being inserted later and at random
[155].
The idea that the genomes of the first organisms contained non-coding sequences can be
Précédent

- 30/799

Suivant