composition. Further regional differences can be
demonstrated in the base composition of the vertebrates. The genome is apparently split into
internally homologous segments of up to 300 kb
that differ in their G+C content (isochores);
these correspond to the microscopically visible
chromosome bands. The differing composition of
the isochores has consequences not only for the
sequences of the coding and non-coding DNA
regions, but also for the transcripts and proteins.
A high proportion of G+C-rich isochores is found
only in birds and mammals; in lower vertebrates
they are seldom found, if they occur at all. From
human DNA, for example, various fractions with
G+C values of between 36.7 and 49.2 % can be
isolated [147,311].
The transfer of methyl groups to single DNA
bases results in the formation of 5-methy1cytosine
(msC) and other methylated bases. In animals
and other eukaryotes, cytosine is methylated only
when in the sequence CG. The resulting mSCG
sequence promotes the methylation of the complementary sequence GC, so that methylated
residues are often found in pairs as mSCG/GmsC
[7]. In contrast to the situation in prokaryotes such as E. coli, some methylatable bases
always remain unmethylated in the eukaryotes.
The degree of methylation is species and tissue
specific and dependent on development. In vertebrates about 3-6 % of cytosine residues are
methylated, whereas in invertebrates the values
are mostly lower; many insects and unicellular
organisms contain no methy1cytosines. In contrast, up to 30 % of the cytosine residues are
methylated in plants. Relatively high degrees of
methylation are to be found in the GC-rich 5'
non-translated (NT) regions of genes [7]. The
DNA of embryos and tumour cells is always
heavily methylated, and the transcriptionally
inactive DNA of spermatozoa is more heavily
methylated than that of active cells [5, 7, 79].
Using immunological methods, other methylated
DNA residues, e.g. 6-methyladenine or 7methylguanine, can be found in molar proportions of several percent of the corresponding
bases; this occurs, e.g., in the citrus scale Planococcus citri or, in somewhat lower proportions, in
Drosophila, human placenta, calf thymus and rat
sperm [3]. 6-Methyladenine, but not 5methy1cytosine, is also present in the macronucleus of several ciliates [35, 159]. Methylation is
apparently of importance for transcription of the
DNA; for this reason the process is presently the
subject of intensive study [5, 79]. However, the
relationship between gene methylation and activ2.1.2 Base Sequence and Gene Structure
11
ity is by no means straightforward; active genes,
as a whole or at particular points, are often hypomethylated, but they may also be completely
methylated, or there may be no correlation
between methylation and gene expression [113].
2.1.2 Base Sequence and Gene Structure
The information contained in the nucleic acids is
encoded in the order of the subunits, the base or
nucleotide sequence. The base sequence in the
DNA of any given organism is the result of
chance mutation and selection during evolution
related to various functional requirements: for
the encoding of polypeptides and RNAs with specific functions; for initiation, termination and
regulation of transcription; for the coiling and
packing of the DNA in the chromosomes; for the
regulation of chromosome pairing during meiosis; and for recombination. In modern biology,
one of the most fascinating possibilities for gaining new insight lies in our ability to read the DNA
text and to trace therein the changes that
occurred during the millions of years of evolution.
Automatic DNA sequencing methods have
developed to such an extent that up to 50000 bp
can be sequenced per day [284], and the large
amount of accumulated data can now be stored
and evaluated only with the computer [459]. The
DNA databanks established in Los Alamos and
Heidelberg contained 6600 sequences in 1985,
amounting to almost 6 million bp, and the problem of coping with the flood of data has since
become acute [254]. Known sequences from the
human genome, ribosomal RNAs, transfer RNAs
and their genes, histones and histone genes, etc.
are now regularly compiled in the journal Nucleic
Acids Research. Approximately 60 % of the
genome of the nematode Caenorhabditis elegans,
which with its 80 million bp is one of the smallest
known eumetazoan genomes, has already been
cloned, as has the genome of baker's yeast, Saccharomyces cerevisiae [94]. Given the present
sequencing capacity, the plan to sequence the
complete 3000 million bp of the human genome
no longer sounds utopian [468]. In 1989, about
6500 human DNA sequences were known, including approximately 1600 protein genes; in the
order of 1000 new sequences are added each year
[494]. DNA sequencing is much more efficient
than protein sequencing and amino acid sequences are today almost entirely derived from the
corresponding DNA sequences. The starting
Précédent

- 26/799

Suivant