42
2 Nucleic Acids and Nuclear Proteins
expression will provide decisive insight into the
processes of cell differentiation and morphogenesis. The control mechanisms of gene expression,
particularly of transcription, are at the centre of
one of the most actively researched areas of modern biology. Investigations in invertebrates, especially Drosophila melanogaster and the nematode
Caenorhabditis elegans, have become particularly
important; the unique suitability of the latter for
such studies is based on the fact that each individual contains the same number of cells and the
developmental fate of each of these cells has been
described.
The most important mechanism for transcription control is the interaction between regulatory
proteins and particular DNA sequences, especially promoters and enhancers. The DNAbinding domains of the regulatory proteins have
four different structural motifs [215, 253, 423]:
the helix-tum-helix motif, which consists of two
a-helices separated by a B turn, was discovered in
gene activation and inhibition in prokaryotes, but
is also found in vertebrates, e.g. in the form of
the determination factors of muscle development
[44]. Related structures are the homeodomains
that are coded by the homeobox, a DNA element
found in many very different eukaryotes. A second DNA-binding motif is the zinc finger, which
was discovered in the transcription factor TFIIIA
of the clawed frog, Xenopus laevis. This contains
nine repetitive sequences of about 30 amino
acids, each with two terminal cysteine and histidine residues that together bind a zinc molecule,
whilst the amino acids in between project as a finger [76]. Many proteins with this motif are now
known to exist in a variety of organisms from
baker's yeast to man; they function as transcription activators or determination factors. In the
vertebrates, the zinc fingers are encoded in multigene families [75, 235]. In Drosophila melanogaster the products of, for example, the segmentation genes "Krtippel" (Kr) and "hunchback" (hb)
are zinc finger proteins, as are the products of the
glass gene that is required for the differentiation
of the photoreceptor cells [309]. The third DNAbinding motif, found in the steroid receptors of
the vertebrates, is also a zinc finger in which,
however, the zinc is bound coordinately by four
cysteine residues. To this class belong, for example, the products of the Drosophila "knirps" (kni)
gene, which is responsible for abdominal segmentation, and the "Seven-up" gene, the absence
of which causes the photoreceptor cells Rl, R3,
R4 and R6 of the eye to assume the characteristics of the R7 type [302, 319]. The fourth motif,
the leucine zipper, contains four or five leucine
residues, each separated by exactly seven amino
acids; as a result, the leucines all lie on the same
side of the a-helix. Interaction between these leucine residues stabilizes a heterodimeric, DNAbinding structure such as is found, for example, in
the products of the oncogenes jun, fos and myc
[215, 423]. Active chromatin regions are always
more sensitive to nucleases; the areas in which
the DNA is free of nucleosomes, and is therefore
easily accessible for regulatory proteins, show a
tenfold higher sensitivity and are termed "hypersensitive sites". These areas are rich in enzymes
such as topoisomerases I and II, RNA polymerase II, and transcription factors [166].
Still to be explained is the mechanism whereby
the binding of a regulatory protein to a DNA
sequence can influence the expression of a gene
that is 100 bp or even 1000 bp removed. Four
mechanisms have been suggested:
1. Two proteins, bound to different sites on the
DNA, interact through the formation of a
DNA loop ("looping").
2. The binding of one protein so alters the conformation of the DNA that the binding of
another protein (e.g. an enzyme) is facilitated
("twisting") .
3. The regulatory protein recognizes a specific
DNA sequence and then moves from there to
another site, at which it initiates transcription,
for example, by interaction with the promoter
("sliding") .
4. Binding of a protein to one site facilitates protein binding at a neighbouring site and so on
until the promoter is reached ("oozing") [352].
Genes that in different cell types are activated at
different times should possess several regulatory
sequences that interact with different factors.
There are various possible patterns of organization:
1. The individual members of a multi-gene family
could carry different control elements; the
best-known example here is the globin gene
family.
2. One gene can possess different control elements for the same initiation signal, as found, for
example, in the genes for the yolk protein, the
heat-shock protein hsp26, and the "white" and
"fushi tarazu" genes of Drosophila.
3. One gene can possess two independently regulated initiation signals, as is seen, for example,
in the a-amylase gene of the mouse and the
gene for alcohol dehydrogenase in Drosophila
melanogaster [137].
2 Nucleic Acids and Nuclear Proteins
expression will provide decisive insight into the
processes of cell differentiation and morphogenesis. The control mechanisms of gene expression,
particularly of transcription, are at the centre of
one of the most actively researched areas of modern biology. Investigations in invertebrates, especially Drosophila melanogaster and the nematode
Caenorhabditis elegans, have become particularly
important; the unique suitability of the latter for
such studies is based on the fact that each individual contains the same number of cells and the
developmental fate of each of these cells has been
described.
The most important mechanism for transcription control is the interaction between regulatory
proteins and particular DNA sequences, especially promoters and enhancers. The DNAbinding domains of the regulatory proteins have
four different structural motifs [215, 253, 423]:
the helix-tum-helix motif, which consists of two
a-helices separated by a B turn, was discovered in
gene activation and inhibition in prokaryotes, but
is also found in vertebrates, e.g. in the form of
the determination factors of muscle development
[44]. Related structures are the homeodomains
that are coded by the homeobox, a DNA element
found in many very different eukaryotes. A second DNA-binding motif is the zinc finger, which
was discovered in the transcription factor TFIIIA
of the clawed frog, Xenopus laevis. This contains
nine repetitive sequences of about 30 amino
acids, each with two terminal cysteine and histidine residues that together bind a zinc molecule,
whilst the amino acids in between project as a finger [76]. Many proteins with this motif are now
known to exist in a variety of organisms from
baker's yeast to man; they function as transcription activators or determination factors. In the
vertebrates, the zinc fingers are encoded in multigene families [75, 235]. In Drosophila melanogaster the products of, for example, the segmentation genes "Krtippel" (Kr) and "hunchback" (hb)
are zinc finger proteins, as are the products of the
glass gene that is required for the differentiation
of the photoreceptor cells [309]. The third DNAbinding motif, found in the steroid receptors of
the vertebrates, is also a zinc finger in which,
however, the zinc is bound coordinately by four
cysteine residues. To this class belong, for example, the products of the Drosophila "knirps" (kni)
gene, which is responsible for abdominal segmentation, and the "Seven-up" gene, the absence
of which causes the photoreceptor cells Rl, R3,
R4 and R6 of the eye to assume the characteristics of the R7 type [302, 319]. The fourth motif,
the leucine zipper, contains four or five leucine
residues, each separated by exactly seven amino
acids; as a result, the leucines all lie on the same
side of the a-helix. Interaction between these leucine residues stabilizes a heterodimeric, DNAbinding structure such as is found, for example, in
the products of the oncogenes jun, fos and myc
[215, 423]. Active chromatin regions are always
more sensitive to nucleases; the areas in which
the DNA is free of nucleosomes, and is therefore
easily accessible for regulatory proteins, show a
tenfold higher sensitivity and are termed "hypersensitive sites". These areas are rich in enzymes
such as topoisomerases I and II, RNA polymerase II, and transcription factors [166].
Still to be explained is the mechanism whereby
the binding of a regulatory protein to a DNA
sequence can influence the expression of a gene
that is 100 bp or even 1000 bp removed. Four
mechanisms have been suggested:
1. Two proteins, bound to different sites on the
DNA, interact through the formation of a
DNA loop ("looping").
2. The binding of one protein so alters the conformation of the DNA that the binding of
another protein (e.g. an enzyme) is facilitated
("twisting") .
3. The regulatory protein recognizes a specific
DNA sequence and then moves from there to
another site, at which it initiates transcription,
for example, by interaction with the promoter
("sliding") .
4. Binding of a protein to one site facilitates protein binding at a neighbouring site and so on
until the promoter is reached ("oozing") [352].
Genes that in different cell types are activated at
different times should possess several regulatory
sequences that interact with different factors.
There are various possible patterns of organization:
1. The individual members of a multi-gene family
could carry different control elements; the
best-known example here is the globin gene
family.
2. One gene can possess different control elements for the same initiation signal, as found, for
example, in the genes for the yolk protein, the
heat-shock protein hsp26, and the "white" and
"fushi tarazu" genes of Drosophila.
3. One gene can possess two independently regulated initiation signals, as is seen, for example,
in the a-amylase gene of the mouse and the
gene for alcohol dehydrogenase in Drosophila
melanogaster [137].
