1.2.2 Protein
Once the genetic code is transcribed into an RNA molecule and the intronic regions are
removed, the mRNA can be translated into a functional protein. The molecular process
from a mRNA to a protein is called translation. Proteins consist of amino acids and are, as
well as DNA and some RNAs, formed 3-dimensionally. There are generally 20 different
amino acids, which can be used to build proteins. An amino acid (AA) chain with less than
40 AAs is called a polypeptide. There are 64 possible permutations of three-letter
sequences that can be made from the four nucleotides. Sixty-one codons represent amino
acids, and three are stop signals. Although each codon is specific for only one amino acid
(or one stop signal), a single amino acid may be coded for by more than one codon. This
characteristic of the genetic code is described as degenerate or redundant. The genetic
codons are illustrated in Table 1.2, representing all nucleotide triplets and their associated
amino acid(s), START or STOP signals, respectively. Just like DNA and RNA, a protein
can also be described by its sequence, however, the 3-dimensional structure based on the
biochemical properties of the amino acids and the milieu, in which proteins are folding, are
much more essential for those building blocks of life.
To keep complex things simple, the way our genes (DNA) are converted to a
transportable messenger system (mRNA) to a functional protein is illustrated in
Fig. 1.4 [9].
The above-mentioned molecular structures can be analyzed in detail by the Next
Generation Sequencing (NGS) technology, enabling a genome wide insight into the
organization and functionality of the genome and all other molecules resulting therefrom
[10]. These will be explained in the following chapters.
1.2.3 Other Important Features of the Genome
Before we can go into detail in terms of sequencing technologies and application, we have
to mention some other important features of the genome and its associated molecules. As
you already know, the human genome consists of ~3 billion base pairs and is harbored in
almost every single cell within the human organism (~10
14 cells). Actually, that is quite a
lot. However, only roughly 1.5% of the whole genome is coding for proteins. What about
the rest? Ninety-eight percent of the human genome useless? Obviously not. The ENCODE
(Encyclopedia of DNA Elements) project has found that 78% of non-coding DNA serve a
defined purpose [11]. Well, think about all the different tasks of all the different cell types
making up a human being. Not every single cell has to be able to do everything—they are
Fig. 1.3 Splice signals usually occur as the first and last dinucleotides of an intron
6
A. Bosserhoff and M. Kappelmann-Fenzl
Once the genetic code is transcribed into an RNA molecule and the intronic regions are
removed, the mRNA can be translated into a functional protein. The molecular process
from a mRNA to a protein is called translation. Proteins consist of amino acids and are, as
well as DNA and some RNAs, formed 3-dimensionally. There are generally 20 different
amino acids, which can be used to build proteins. An amino acid (AA) chain with less than
40 AAs is called a polypeptide. There are 64 possible permutations of three-letter
sequences that can be made from the four nucleotides. Sixty-one codons represent amino
acids, and three are stop signals. Although each codon is specific for only one amino acid
(or one stop signal), a single amino acid may be coded for by more than one codon. This
characteristic of the genetic code is described as degenerate or redundant. The genetic
codons are illustrated in Table 1.2, representing all nucleotide triplets and their associated
amino acid(s), START or STOP signals, respectively. Just like DNA and RNA, a protein
can also be described by its sequence, however, the 3-dimensional structure based on the
biochemical properties of the amino acids and the milieu, in which proteins are folding, are
much more essential for those building blocks of life.
To keep complex things simple, the way our genes (DNA) are converted to a
transportable messenger system (mRNA) to a functional protein is illustrated in
Fig. 1.4 [9].
The above-mentioned molecular structures can be analyzed in detail by the Next
Generation Sequencing (NGS) technology, enabling a genome wide insight into the
organization and functionality of the genome and all other molecules resulting therefrom
[10]. These will be explained in the following chapters.
1.2.3 Other Important Features of the Genome
Before we can go into detail in terms of sequencing technologies and application, we have
to mention some other important features of the genome and its associated molecules. As
you already know, the human genome consists of ~3 billion base pairs and is harbored in
almost every single cell within the human organism (~10
14 cells). Actually, that is quite a
lot. However, only roughly 1.5% of the whole genome is coding for proteins. What about
the rest? Ninety-eight percent of the human genome useless? Obviously not. The ENCODE
(Encyclopedia of DNA Elements) project has found that 78% of non-coding DNA serve a
defined purpose [11]. Well, think about all the different tasks of all the different cell types
making up a human being. Not every single cell has to be able to do everything—they are
Fig. 1.3 Splice signals usually occur as the first and last dinucleotides of an intron
6
A. Bosserhoff and M. Kappelmann-Fenzl
