76
cytosine methylation, histone acetylation, and changes in
chromatin structure may lead to a subsequently altered transcriptome (Alberts et al. 2008).
Due to the translation of mRNA into amino acids via the
triplet code, proteins are in a qualitative sense direct product
of genes with mRNA transcripts as intermediates. This
allows functional predictions of genes via comparison of
sequence similarities to annotated genes in a highly curated
database, such as NCBI RefSeq (O’Leary et al. 2016).
In eukaryotes, the RNA sequence is, nevertheless, subject
to possible modifications, which may impede the recognition
of a gene-protein pair. Variable intron removal from maturing mRNAs by splicing may lead to multiple isoforms from
a single pre-mRNA (Alberts et al. 2008). Further, RNA editing (see example in section “Response to environmental
cues”) may introduce sequence alterations as a co- or posttranscriptional modification, not to be confused with decapping, splicing, and poly(A)-removal (see e.g., Klug et al.
2012; Liew et al. 2017).
Sequence Alterations Influence Protein
Functioning
Non-synonymous sequence alterations, i.e. single nucleotide exchanges, deletions, or insertions, may significantly
influence or disrupt protein functioning. Firstly, a protein’s
physiological role is sensitive to secondary and tertiary
structure formation and stability (e.g., α-helix and cysteine
double bounds, respectively), which may be significantly
altered due to aforementioned non-synonymous sequence
alterations. Secondly, the phosphorylation of serine, threonine, and tyrosine, as well as acetylation and ubiquitylation of lysine are major post-translational modifications,
which are involved in triggering activation and degradation (reviewed in Klug et al. 2012; Ruggles et al. 2017).
Thus, sequence alterations which lead to the exchange of
one of these four amino acids are likely to affect the protein’s performance. Lastly, guiding and localization
sequences are essential to position proteins in cellular
compartments or membranes. For example, the nuclear
membrane of most eukaryotic cells is freely permeable to
molecules up to 9 nm. Macromolecules of greater sizes
depend on a specific nuclear localization sequence (NLS),
which mediates the transport. Alteration of a single amino
acid may result in a dysfunction of the NLS and the
decreased transport efficiency of the macromolecule into
the nucleus (Zanta et al. 1999).
Consequently, complex reactions such as protein-protein
interactions, transcription cascades, signaling networks, and
metabolic pathways may be altered by single nucleotide
exchanges (Kim et al. 2016).
Quantitative Regulation of the Proteome
The physiological roles of RNA reach far beyond the gene to
protein transmission, where (pre)-mRNA, rRNA, and tRNA
are allocated. For instance, the translation-regulatory roles of
miRNAs have been discovered in 1993 (Almeida et al. 2011;
see section “Functionality”). In humans, for example, at least
70% of the genome is transcribed into RNA, but only about
2% are effectively translated into protein (Pheasant and
Mattick 2007). Consequently, immense proportions of the
genome are suggested to encode for quantitative regulation,
which can be detected with current omics approaches (Klug
et al. 2012). The current state of knowledge considers the
abundance of mRNA transcripts to explain up to 84% of the
respective protein concentration. This value may vary
depending on the respective mRNA, mainly attributable to
sequence- or splice isoform-dependent translation rates (Liu
et al. 2016). Additionally, induced changes in gene expression, e.g., due to environmental cues, may only be detectable
in the proteome after a lag phase (e.g., 6–7 h in mammals;
see also section “Response to environmental cues”).
The number of copies per gene does not generally define
respective transcript nor protein abundances. Genetic diseases or tumors may induce gene copy number alterations
(CNAs). In such cases, transcriptome and proteome do
mostly not exhibit the same fold changes as could be expected
from the CNAs in the genome. Negative feedback loops,
called buffering, may occur on the transcriptional and translational level. There are, however, plenty of sequencespecific exceptions to this general pattern, which are,
therefore, possibly involved in the symptomatic (Liu et al.
2016 and references therein).
Metabolomics
The entirety of small molecules within an organism, the
metabolome, constitutes a biochemical representation. It is
substance to continuous turn-over, alteration, and relocation
by the physiological machinery of RNAs and, most of all,
proteins (e.g., Patti et al. 2012; Beale et al. 2016). While targeted metabolomics assesses only a fraction of particular
interest, newly emerged technologies enable untargeted
detection and quantification of almost the entire metabolome
(Patti et al. 2012).
Untargeted metabolomics combined with genomic and/or
transcriptomic data may allow the inference of gene and protein function, as well as metabolic cascades and pathways. It
becomes possible to detect physiological attributes such as
the use of substrates, secondary metabolite secretion, or possible inter-individual signaling, and connect these to the
presence or expression of genes (Freilich et al. 2011;
J. D. Brüwer and H. Buck-Wiese
cytosine methylation, histone acetylation, and changes in
chromatin structure may lead to a subsequently altered transcriptome (Alberts et al. 2008).
Due to the translation of mRNA into amino acids via the
triplet code, proteins are in a qualitative sense direct product
of genes with mRNA transcripts as intermediates. This
allows functional predictions of genes via comparison of
sequence similarities to annotated genes in a highly curated
database, such as NCBI RefSeq (O’Leary et al. 2016).
In eukaryotes, the RNA sequence is, nevertheless, subject
to possible modifications, which may impede the recognition
of a gene-protein pair. Variable intron removal from maturing mRNAs by splicing may lead to multiple isoforms from
a single pre-mRNA (Alberts et al. 2008). Further, RNA editing (see example in section “Response to environmental
cues”) may introduce sequence alterations as a co- or posttranscriptional modification, not to be confused with decapping, splicing, and poly(A)-removal (see e.g., Klug et al.
2012; Liew et al. 2017).
Sequence Alterations Influence Protein
Functioning
Non-synonymous sequence alterations, i.e. single nucleotide exchanges, deletions, or insertions, may significantly
influence or disrupt protein functioning. Firstly, a protein’s
physiological role is sensitive to secondary and tertiary
structure formation and stability (e.g., α-helix and cysteine
double bounds, respectively), which may be significantly
altered due to aforementioned non-synonymous sequence
alterations. Secondly, the phosphorylation of serine, threonine, and tyrosine, as well as acetylation and ubiquitylation of lysine are major post-translational modifications,
which are involved in triggering activation and degradation (reviewed in Klug et al. 2012; Ruggles et al. 2017).
Thus, sequence alterations which lead to the exchange of
one of these four amino acids are likely to affect the protein’s performance. Lastly, guiding and localization
sequences are essential to position proteins in cellular
compartments or membranes. For example, the nuclear
membrane of most eukaryotic cells is freely permeable to
molecules up to 9 nm. Macromolecules of greater sizes
depend on a specific nuclear localization sequence (NLS),
which mediates the transport. Alteration of a single amino
acid may result in a dysfunction of the NLS and the
decreased transport efficiency of the macromolecule into
the nucleus (Zanta et al. 1999).
Consequently, complex reactions such as protein-protein
interactions, transcription cascades, signaling networks, and
metabolic pathways may be altered by single nucleotide
exchanges (Kim et al. 2016).
Quantitative Regulation of the Proteome
The physiological roles of RNA reach far beyond the gene to
protein transmission, where (pre)-mRNA, rRNA, and tRNA
are allocated. For instance, the translation-regulatory roles of
miRNAs have been discovered in 1993 (Almeida et al. 2011;
see section “Functionality”). In humans, for example, at least
70% of the genome is transcribed into RNA, but only about
2% are effectively translated into protein (Pheasant and
Mattick 2007). Consequently, immense proportions of the
genome are suggested to encode for quantitative regulation,
which can be detected with current omics approaches (Klug
et al. 2012). The current state of knowledge considers the
abundance of mRNA transcripts to explain up to 84% of the
respective protein concentration. This value may vary
depending on the respective mRNA, mainly attributable to
sequence- or splice isoform-dependent translation rates (Liu
et al. 2016). Additionally, induced changes in gene expression, e.g., due to environmental cues, may only be detectable
in the proteome after a lag phase (e.g., 6–7 h in mammals;
see also section “Response to environmental cues”).
The number of copies per gene does not generally define
respective transcript nor protein abundances. Genetic diseases or tumors may induce gene copy number alterations
(CNAs). In such cases, transcriptome and proteome do
mostly not exhibit the same fold changes as could be expected
from the CNAs in the genome. Negative feedback loops,
called buffering, may occur on the transcriptional and translational level. There are, however, plenty of sequencespecific exceptions to this general pattern, which are,
therefore, possibly involved in the symptomatic (Liu et al.
2016 and references therein).
Metabolomics
The entirety of small molecules within an organism, the
metabolome, constitutes a biochemical representation. It is
substance to continuous turn-over, alteration, and relocation
by the physiological machinery of RNAs and, most of all,
proteins (e.g., Patti et al. 2012; Beale et al. 2016). While targeted metabolomics assesses only a fraction of particular
interest, newly emerged technologies enable untargeted
detection and quantification of almost the entire metabolome
(Patti et al. 2012).
Untargeted metabolomics combined with genomic and/or
transcriptomic data may allow the inference of gene and protein function, as well as metabolic cascades and pathways. It
becomes possible to detect physiological attributes such as
the use of substrates, secondary metabolite secretion, or possible inter-individual signaling, and connect these to the
presence or expression of genes (Freilich et al. 2011;
J. D. Brüwer and H. Buck-Wiese
