12
2 Nucleic Acids and Nuclear Proteins
sequence may be that of either a DNA complementing an RNA sequence (cDNA) or genomic
DNA; the amounts of DNA required for
sequencing are obtained through cloning reactions [30,204]. The multiplication of complete
genes from genomic DNA is now possible using
the polymerase chain reaction (peR), i. e. a
DNA-polymerase-catalysed DNA synthesis using
a large excess of short DNA primers that contain
a partial sequence for the gene of interest
[39,267].
The gene concept has undergone many
changes as our knowledge of mode of gene action
has increased. After the classical definition of the
gene as a unit of mutation and recombination,
and then the "one gene, one polypeptide" principle, came the description in the 1960s of a gene as
a stretch of DNA that is transcribed as a unit and
that codes for a functional RNA (messenger
RNA, ribosomal RNA, transfer RNA) [168].
Although, in the meantime, in animal genomes
phenomena have been observed which do not fit
this definition, e.g. overlapping and superimposed genes or the formation of multiple RNAs
from a single gene, there is still much to be said
for retaining the definition of a gene as a transcription unit (Fig. 2.1). The transcribed region of
the DNA is always longer than the sequence coding for the finished product; in addition, in front
of ("upstream", 5' end) or behind ("downstream", 3' end) the transcribed region lie DNA
sequences of importance as signals for the initiation and termination of transcription. The question thus arises of how the gene as a functional
unit is to be defined on the DNA. For this, one
can exploit the changes in chromatin structure,
shown by active genes, that make them more susceptible to DNase I digestion. Thus, the fibroin
gene of the silkworm Bornbyx rnori is very sensitive to DNase when it is actively transcribed in the
distal part of the silk gland of fifth-stage larvae,
but is essentially insensitive to DNase I in its
inactive state in the middle part of the silk gland
of such larvae, or in later developmental stages
[236]. In the chicken, the DNase-I-sensitive DNA
region of the ovalbumin gene and the closely
related genes "X" and "Y" has a length of 100 kb,
that of the lysozyme gene is 24 kb, and for
glyceraldehyde-3-phosphate dehydrogenase it is
12 kb; in all cases, the boundaries at the 3' end
are sharp and those at the 5' end are more
variable [213]. The signals for initiation and termination of transcription have different positions
and structures according to the gene product and
the relevant RNA polymerase, i. e. the Pol I transcribed ribosomal RNA genes, the Pol II transcribed protein-coding genes, and the Pol III
transcribed genes for small RNAs. The transcription signals will be discussed in connection with
the process of transcription. The start and finish
of translation are also marked by signals on the
DNA and mRNA, respectively. Translation begins at the first AUG of the mRNA and ends at
the first UAA, UAG or UGA stop codon.
Normally the genes of the eukaryote genome
are arranged in a linear fashion. In prokaryotes,
on the other hand, overlapping genes are not
uncommon; these lie either on different DNA
strands and are read in opposite directions, or on
the same strand and have interlocking reading
frames. This situation, although it allows a higher
density of information on the DNA and is advantageous for small genomes, hinders the evolution
of DNA sequences. It is also particularly surprising to find such phenomena in animals that have
a large surplus of DNA. The DNA flanking the
protamine gene of the rainbow trout, Salrno
gairdneri, includes an extra TAT A box in front of
the promoter and the region is transcribed into
two different functional mRNAs. One of the
mRNAs encodes protamine, and the other, which
arises from a shifted reading frame and is found in
large amounts only in the brain, produces the
proline-rich Y protein (Fig. 2.2). The function of
the Y protein is not known; it is significantly
homologous to a proline-rich phosphoprotein
from human saliva, the avian sarcoma virus
(ASV) protein P19, and the product of the myc
oncogene. The Y gene breaks the "first AUG
rule" in so far as translation begins only at the
third ATG after the TATA box. Interestingly, the
protamine gene is flanked by long terminal
repeats (LTRs) that are characteristic of ASV and
other viruses. There is also a region in the ASV
genome that codes for three different proteins
using shifted reading frames, and it is easy to
imagine that the anomalies in animals originated
with the insertion of a viral genome [212]. At the
Initiation
Poly (Al
Fig. 2.1. A plan of the transcription unit of a proteincoding gene trancribed by
polymerase II
Control
I
I
Termination
region
+ Exons
+
region
- - - - - - - - TATAAT ~ A A U A A A - - - - - - -
2 Nucleic Acids and Nuclear Proteins
sequence may be that of either a DNA complementing an RNA sequence (cDNA) or genomic
DNA; the amounts of DNA required for
sequencing are obtained through cloning reactions [30,204]. The multiplication of complete
genes from genomic DNA is now possible using
the polymerase chain reaction (peR), i. e. a
DNA-polymerase-catalysed DNA synthesis using
a large excess of short DNA primers that contain
a partial sequence for the gene of interest
[39,267].
The gene concept has undergone many
changes as our knowledge of mode of gene action
has increased. After the classical definition of the
gene as a unit of mutation and recombination,
and then the "one gene, one polypeptide" principle, came the description in the 1960s of a gene as
a stretch of DNA that is transcribed as a unit and
that codes for a functional RNA (messenger
RNA, ribosomal RNA, transfer RNA) [168].
Although, in the meantime, in animal genomes
phenomena have been observed which do not fit
this definition, e.g. overlapping and superimposed genes or the formation of multiple RNAs
from a single gene, there is still much to be said
for retaining the definition of a gene as a transcription unit (Fig. 2.1). The transcribed region of
the DNA is always longer than the sequence coding for the finished product; in addition, in front
of ("upstream", 5' end) or behind ("downstream", 3' end) the transcribed region lie DNA
sequences of importance as signals for the initiation and termination of transcription. The question thus arises of how the gene as a functional
unit is to be defined on the DNA. For this, one
can exploit the changes in chromatin structure,
shown by active genes, that make them more susceptible to DNase I digestion. Thus, the fibroin
gene of the silkworm Bornbyx rnori is very sensitive to DNase when it is actively transcribed in the
distal part of the silk gland of fifth-stage larvae,
but is essentially insensitive to DNase I in its
inactive state in the middle part of the silk gland
of such larvae, or in later developmental stages
[236]. In the chicken, the DNase-I-sensitive DNA
region of the ovalbumin gene and the closely
related genes "X" and "Y" has a length of 100 kb,
that of the lysozyme gene is 24 kb, and for
glyceraldehyde-3-phosphate dehydrogenase it is
12 kb; in all cases, the boundaries at the 3' end
are sharp and those at the 5' end are more
variable [213]. The signals for initiation and termination of transcription have different positions
and structures according to the gene product and
the relevant RNA polymerase, i. e. the Pol I transcribed ribosomal RNA genes, the Pol II transcribed protein-coding genes, and the Pol III
transcribed genes for small RNAs. The transcription signals will be discussed in connection with
the process of transcription. The start and finish
of translation are also marked by signals on the
DNA and mRNA, respectively. Translation begins at the first AUG of the mRNA and ends at
the first UAA, UAG or UGA stop codon.
Normally the genes of the eukaryote genome
are arranged in a linear fashion. In prokaryotes,
on the other hand, overlapping genes are not
uncommon; these lie either on different DNA
strands and are read in opposite directions, or on
the same strand and have interlocking reading
frames. This situation, although it allows a higher
density of information on the DNA and is advantageous for small genomes, hinders the evolution
of DNA sequences. It is also particularly surprising to find such phenomena in animals that have
a large surplus of DNA. The DNA flanking the
protamine gene of the rainbow trout, Salrno
gairdneri, includes an extra TAT A box in front of
the promoter and the region is transcribed into
two different functional mRNAs. One of the
mRNAs encodes protamine, and the other, which
arises from a shifted reading frame and is found in
large amounts only in the brain, produces the
proline-rich Y protein (Fig. 2.2). The function of
the Y protein is not known; it is significantly
homologous to a proline-rich phosphoprotein
from human saliva, the avian sarcoma virus
(ASV) protein P19, and the product of the myc
oncogene. The Y gene breaks the "first AUG
rule" in so far as translation begins only at the
third ATG after the TATA box. Interestingly, the
protamine gene is flanked by long terminal
repeats (LTRs) that are characteristic of ASV and
other viruses. There is also a region in the ASV
genome that codes for three different proteins
using shifted reading frames, and it is easy to
imagine that the anomalies in animals originated
with the insertion of a viral genome [212]. At the
Initiation
Poly (Al
Fig. 2.1. A plan of the transcription unit of a proteincoding gene trancribed by
polymerase II
Control
I
I
Termination
region
+ Exons
+
region
- - - - - - - - TATAAT ~ A A U A A A - - - - - - -
